Agentic Research

Local vs cloud (run it yourself or call an API)

Also: 本地部署 · 雲端 API · self-hosted · 本地跑還是用 API · on-premise

Whether the model runs on your own machine (local) or you call someone else's over the network (cloud API). The real difference is not "free vs paid" but several axes: privacy, cost at volume, latency, capability ceiling and maintenance burden.

When you will meet it

This is the first architecture decision every AI project must make, sooner or later. People who reduce it to "local is free, cloud costs money" get burned in both directions: local ignores hardware and ops cost, cloud ignores data exposure and the bill once you scale. Knowing the axes tells you what you are trading for what.

An analogy

Like buying a car versus taking taxis. Owning: expensive up front, needs maintenance and parking, but you can drive anytime, your things stay in the car (privacy), and heavy long-term use actually saves money. Taxis: no upkeep, hop in and go (a high capability ceiling, on demand), but every trip costs, the trip log is in someone else's hands, and at peak you may not get one or may pay more.

Minimal example

別用「免費 vs 付費」思考,用這幾條軸:
  隱私       敏感資料能不能離開你的機器?
  規模成本   用量越大,雲端每 token 帳單越可觀;本地硬體是一次性
  延遲       本地沒有網路往返;雲端要加上上行與服務端排隊
  能力上限   最大的模型通常只在雲端;本地受你的 VRAM/記憶體限制
  維運負擔   本地要你自己顧當機、更新、監控;雲端是別人的問題

For most teams the right answer is not either/or but a mix: sensitive or high-frequency work goes local, work needing the strongest capability or burst volume goes cloud. Decide where to draw the line only after thinking through what each axis demands.

What people get wrong

  • Assuming "local = free". Hardware depreciation, power and your ops time are all costs; at high volume cloud can actually be cheaper.
  • Assuming "cloud = data is safe". Your prompts and data travel to someone else's machine; for sensitive content this can be a compliance red line, which is a different question from security.
  • Treating latency as one thing. Local latency comes from hardware; cloud latency adds network round-trips and server-side queuing. The bottlenecks are entirely different.

Related terms

Next