Agentic Research

Harness Engineering: The Workbench Around Your Agent

2026/09/2912 min readBryan Chan閱讀中文原文
TopicsHarness EngineeringAgentAI App DevelopmentPrompt EngineeringMCP

The Bottom Line First

You built an Agent Demo for your team, your boss was very satisfied after seeing it, and told you to push it to production. But as soon as it went live, the Agent changed files it should not have changed, claimed the task was complete while all tests were red, and started talking nonsense after the context was stuffed full.

The problem is not that the model is not smart enough. The problem is that you only installed a brain for it and did not build a workbench for it.

Harness is the runtime system layer wrapped around the AI model. It is responsible for turning vague user goals into executable tasks, feeding the correct context to the model, turning model output into tool calls, and then using tests, permissions, logs, and review mechanisms to determine whether the task is truly complete.

By analogy: the model is a chef's brainpower, while Harness is the entire kitchen. The counter, stove, fire extinguisher, recipe cards, and the sous chef at the pass who checks every dish are all included. No matter how skilled the chef, without a kitchen they can only set up a stall on the roadside.

Harness is the working foundation between the model and the real engineering environment

1. Four Failure Scenarios You Have Definitely Encountered

Moving an Agent from Demo to production usually means hitting these four walls.

Requirements drift. At first you say, "Build an expense tracking tool," and the Agent quickly produces a beautiful page. Then you say, "Add category statistics," and it adds charts. Then you say, "Support multi-user collaboration," and it starts recklessly changing the data structure. By the end, even you cannot clearly explain what the project's core goal is. It is like a construction site with no blueprint: the workers are fast, but the wall is built increasingly crooked.

Context blows up. When an Agent performs complex tasks, it reads many files, runs many commands, and generates many logs. No matter how large the context window is, it is not infinite. The more stuffed it becomes, the more likely the model is to forget early constraints or treat outdated information as the latest fact. It is like an office desk with documents piled as high as a mountain; finding one sheet of paper takes half an hour.

Runnable does not mean correct. A function running on one sample does not mean it satisfies all business edge cases. Monetary calculations use float, rounding rules are wrong, and boundary value handling has holes. Without tests, an Agent can easily write code that "looks right." It is like a bridge that looks stable, but collapses as soon as the first heavy truck drives over it.

Agents are very good at "confidently declaring completion." Many Agents report "completed" without actually verifying. Human engineers at least know whether they ran tests, but if an Agent is not asked to run tests, it may judge only by the surface of the code. It is like a construction crew self-inspecting and self-signing, with no one doing acceptance inspection.

Four walls from Demo to production

2. Why the Old Approach Is Not Enough

Over the past few years, the methodology for AI application development has gone through three upgrades, each solving the gap left by the previous generation.

Phase 1: Prompt Engineering. The core question was, "How do I ask AI so that it answers better?" You write the requirements clearly, and the model generates an answer in one pass. This is sufficient for simple tasks, but it fails with complex engineering because a single sentence cannot contain all the rules of an entire project.

Phase 2: Context Engineering. The core question became, "What context do I give AI so that it works within the correct information?" You feed the model the project directory, existing code, and coding standards. This is much better than writing prompts alone, but it still assumes that the model can get it right in one pass after reading the materials, with no mechanism for self-correction after mistakes.

Phase 3: Harness Engineering. The core question moves one step further: "What kind of task environment, tool boundaries, verification mechanisms, state management, and feedback loops do I give AI so that it can complete complex tasks continuously, reliably, and in an auditable way?"

The difference in levels among the three is clearest with a kitchen metaphor. Prompt Engineering is telling a chef, "Make a dish of braised pork." Context Engineering is placing the recipe, ingredients, and cookware list in front of him. Harness Engineering is building the entire kitchen, specifying which stoves he may use, requiring the temperature to be tested before each dish leaves the pan, and requiring failures to be redone rather than served.

The level differences among the three generations of methodology

3. Core Concepts Breakdown: Nine Modules

A mature Harness can be broken down into nine modules. You do not need to build all of them at once, but you need to know what problem each module solves.

Task specification. It is not a one-line request to "build a feature," but a contract: what the goal is, what is out of scope, what the inputs are, and what the completion criteria are. Turn "a feeling" into "black and white." Just as before building a house you need architectural drawings, clearly stating how many rooms and living rooms there are, and whether there is a balcony.

Context selection. More context is not always better. If you give too much, the model may fail to grasp the key points. The right approach is a short entry point, layered knowledge, and on-demand retrieval. It is like a recipe rack in the kitchen: you do not need to put every recipe in the world on it, only the ones you will use today.

Tool access. Fewer, well-chosen tools matter more than many tools. What an Agent fears most is not having no tools, but having two hundred tools in front of it and not knowing which one to choose. The tools you give it should be like each blade on a Swiss Army knife: each does one thing, and they can be combined.

Project memory. Conversations end, context gets compressed, but files in the repository remain. Important facts should settle into documents rather than staying in the chat history. It is like a bulletin board in the office: post important matters on it, and a newcomer can see them at a glance.

State management. Long tasks cannot rely only on conversation records. Have the Agent write a plan file that records progress, risks, and next steps. Even if the context is compressed, it can return to this file and continue working.

Verification mechanism. This is the soul of the Harness. Unit tests, integration tests, lint, typecheck, build. For an Agent, the most valuable feedback is not "you wrote it well," but "Expected Decimal('85.50'), got Decimal('90.25')". This kind of feedback is very clear, and the Agent can continue fixing based on it.

Permissions and sandboxing. An Agent needs freedom, but not unlimited freedom. Changing tests can be automated, changing core payment logic requires review, deleting data must be refused, and accessing production keys is not allowed. It is like a hospital operating room: the attending physician can operate, while the intern can only watch from the side.

Observability. An Agent failing is not scary; what is scary is that after it fails, you do not know why it failed. Every step should have logs: what files were read, what tools were called, which step failed, and how it was fixed.

Human takeover. Being advanced is not about full automation; it is about automating what should be automated and stopping what should be stopped. The goal of a Harness is not to make people disappear, but to move people from "hand-writing every line of code" to "designing rules, reviewing results, and handling key decisions."

The nine modules of a harness

4. Implementation Checklist: Four Steps to Get Started

You do not need to build a large platform right away. Four steps are enough to get started.

Step 1: Add minimal project context. Add three files: AGENTS.md (project map), docs/architecture.md (module boundaries), docs/testing.md (test commands and strategy). AGENTS.md should only serve as a map, not a manual, and stay under thirty lines. This is like preparing an onboarding guide for a new colleague: it does not need to be exhaustive, but it should keep them from getting lost on their first day.

A sufficient AGENTS.md looks roughly like this:

# Project Map

## What this project does
Explain the core goal in one sentence, no more than two lines.

## How to read the directory
- src/api/     External interfaces
- src/core/    Domain logic, calling external APIs here is prohibited
- tests/       Tests, only add, never remove

## Must read before starting work
- docs/architecture.md   Module boundaries and dependency direction
- docs/testing.md        Test commands and strategy

## Completion criteria (all must pass)
pytest -q
npm run lint
npm run typecheck

## Prohibitions
- Do not directly modify the main branch
- Do not touch existing files under migrations/
- Do not write keys into any code or configuration files

Please note that it has only three key points: what the project does, how to read the directory, and what counts as done. It is intentionally not written as a manual, because the longer the rules are, the easier they are to ignore; what truly matters must be short enough to read in one sitting. The greatest value of this map is that it lets the Agent know where the boundaries are before taking action for the first time, rather than coming back to ask after hitting a wall.

Step 2: Make the "definition of done" machine-checkable. Do not write "code quality should be good"; write pytest -q, npm run lint, npm run typecheck. If a machine can determine it, do not rely on feeling to determine it. This is like a physical exam report: a doctor will not say "your health condition is pretty good"; instead, they give you specific numbers and metrics.

By the same logic, tool boundaries should also be written as a whitelist rather than a description. Below is a minimal viable example of permission rules (assuming it is placed in docs/testing.md):


Can be executed automatically (no approval needed)
  pytest -q
  npm run lint / typecheck / build
  Read any file within the project

Requires manual approval
  Modify any file under src/core/
  Add or change database migrations
  Modify CI configuration and deployment scripts

Always deny
  Write or read production environment keys
  Delete existing tests under tests/
  Execute commands such as rm -rf / DROP TABLE

Note that this checklist uses an "action allowlist" rather than a reminder like "please be careful." Reminders get forgotten, while allowlists are enforced by machines.

Step 3: Agent changes uniformly go through PR. Do not let Agents directly modify the main branch. One branch per task, the Agent submits a diff, CI runs, humans review, and it is merged. The PR template mandates filling in verification results and risk level. This is like dual control at a bank: one person cannot complete a large transaction alone, and a second person must review it.

Step 4: Record failures and improve the Harness. Every time an Agent makes an error, do not just blame the model. Ask: Was the task not written clearly? Was context missing? Are there too many tools? Is there no testing? Are permissions too broad? Then distill the lessons into AGENTS.md, new tests, new lint rules, and new approval rules. This is the closed loop.

Four-step implementation roadmap

5. The Relationship Between MCP and Harness

MCP (Model Context Protocol) is a standard interface for AI applications to connect to external tools and data sources. If we compare an Agent to a chef, MCP is like those standardized plugs in a kitchen: gas stoves, faucets, and range hoods all have uniform specifications, and they can still be used if you switch brands.

But what MCP solves is the problem of "how tools are connected," while Harness Engineering solves the problem of "how tools are used safely, reliably, and verifiably after they are connected." These are two different things.

You can connect a hundred MCP tools to an Agent, but without permissions, verification, and task boundaries, it will still get lost. It is like a kitchen filled with all kinds of high-end appliances, but with no recipes, no serving process, and no sous-chef inspection, the chef will still be flustered.

MCP is the standardized plug, and Harness is the electrical system of the entire building. No matter how many plugs there are, without fuses and a distribution panel, it will still trip.

6. Three Common Misconceptions

Misconception 1: Longer prompts are better. Not true. Prompts that are too long dilute the key points. The right approach is a short entry point, layered knowledge, clear rules, and letting the Agent read on demand. Just as a good office does not post all rules and regulations on the same wall, but stores them by category and looks them up when needed.

Misconception 2: More tools mean more powerful. Not true. More tools increase the cost of choice. The Agent may turn a low-risk task into a high-risk operation. Tools should be few, stable, and composable. Just as a good kitchen is not filled with two hundred utensils; the chef commonly uses just those few knives and a few pots.

Misconception 3: Harness Engineering is only for big companies. Not true. Small teams need it even more, because small teams do not have as much manual review. Even just adding one AGENTS.md and test commands is much more stable than pure Vibe Coding.

Three Sentences to Take Away

  1. The model is the brain, and Harness is the body. A brain without a workbench can only talk idly in chat; with tools, validation, permissions, and logs, the brain can truly get hands-on and do things.
  2. Completion criteria should not be decided by the model. Tests pass, lint passes, human review passes: only then is it truly complete. "I think it is right" does not count.
  3. Starting with one AGENTS.md is enough. There is no need to build all nine modules at once. First let the Agent have a map, validation commands, and a PR process, and it is already ahead of most teams.

下一步

If you want to keep going further, you can first fill in the foundations at the framework layer (Complete LangChain Tutorial 2026), then see how tools are standardized and integrated (MCP Protocol Explained), and finally tie it together from an architecture perspective (Three-Layer Agent Collaboration Framework).