Agentic Research

Harness

Also: agent harness · 執行框架 · harness 是什麼

The layer of code wrapped around a model: assembling context, calling the model, parsing output, running tools, gating permissions, feeding results back.

When you will meet it

The same model in different harnesses performs very differently. When you choose a tool, you are largely choosing a harness.

An analogy

A horse and its harness. The horse has power, but without the harness that power cannot pull a cart. The model is the horse; the name is literal.

Minimal example

harness 負責的事(模型一件都不做):
  · 決定這次要送哪些上下文進去
  · 把可用工具清單描述給模型
  · 解析模型回傳,判斷是文字還是工具請求
  · 真的去執行工具,並在受限環境裡執行
  · 把結果接回上下文,再問一次
  · 決定什麼時候停

So "the AI did something wrong" usually means a link in the harness failed, not that the model is broken.

How a harness runsFlow diagram: a user request enters the harness, which assembles context (system prompt, tool list, memory), calls the LLM for one forward pass, and parses the output. Plain text converges straight to a final answer; a tool call first passes a permission gate (allow, ask, or deny), runs in a sandbox, and the observation is folded back into the context for another model call. The loop repeats until the model stops requesting tools.HARNESStool callplain textloop backUser requesta task, in plain wordsAssemble contextsystem prompt + tools + memoryCall the LLMone forward passParse outputtext, or a tool call?Permission gateallow / ask / denySandboxed runtool runs constrainedObservationresult folded back inFinal answerno more tool calls
1/8User request
The task arrives in plain language. The harness has to turn it into something a model can act on.
Step 1 of 8 User request

What people get wrong

  • Confusing harness with framework. A framework (like LangChain) is a parts box for building a harness; the harness is the running thing you built.
  • Assuming a new tool means a new model. Many tools share the same models; what differs is the harness.

Related terms

Next