Agentic Research

The Ultimate Comparison of 7 AI Agent Loop Engineering Architectures: From while(true) to Multi-Agent Orchestration

2026/06/1472 min readBryan Chan閱讀中文原文
TopicsLoop EngineeringAgent ArchitectureClaude CodeCursorAider

Core proposition: Every AI Agent is ultimately a loop. But how the loop is designed determines whether the agent is a reliable engineer or a flashy demo. This article compares the loop architectures of seven AI Agents, from source-level dissection to architectural decision analysis. Systems covered: Claude Code · Cursor · Aider · Cline · SWE-agent · OpenHands · OpenClaw Loop Engineering Methodology: Source-level analysis + architecture document comparison + production environment field verification


1. Why the Loop Is Everything for an Agent

In 2026, the AI Agent field reached two points of consensus:

Consensus one: Agent = LLM + Tools + Loop. Without a loop, a model is just a one-shot question-and-answer machine. With a loop, a model can observe the world, take action, receive feedback, and iterate.

Consensus two: The design of the loop matters more than the model itself. A well-designed loop can let a mid-tier model outperform a top-tier model running inside a poor loop.

Vincent Qiao, who analyzed the Claude Code source code, wrote:

"The agent loop is 1,730 lines of a single while(true) that does everything. It streams model responses, executes tools concurrently, compresses context through four layers, recovers from five categories of errors, tracks token budgets with diminishing returns detection, runs stop hooks that can force the model back to work, manages prefetch pipelines for memory and skills, and produces a typed discriminated union of exactly why it stopped."

This while(true) is the shared core of all agent systems, but the way each system implements it differs fundamentally. This article dissects seven major approaches.


2. In-Depth Comparison of Seven Loop Architectures

2.1 Claude Code: One Giant while(true)

Loop type: AsyncGenerator based single-loop

Core structure:

while (true) {
  1. Prepare messages (compress, trim, collapse: five-layer compression pipeline)
  2. Call Claude API (streaming)
  3. Collect model response
  4. Has tool_use? → Execute tools concurrently → Collect results → continue
  5. No tool_use? → Check Stop Hooks → Check recovery needed → Otherwise exit
  6. Assemble [original + assistant response + tool results]
  7. Enter next iteration
}

Key design decisions:

DesignImplementationImpact
Single code pathquery.ts, a single 1,730-line query() function; all entry points (REPL, SDK, subagent, headless) go through itAny improvement automatically benefits all scenarios
Concurrent tool executionStreamingToolExecutor, tools start executing while the model is still emitting outputSignificantly lower latency
Stop HookAn external script forcibly checks the result before the turn ends. If it does not pass, the turn is blocked (up to 8 times)Separates execution from verification
7 recovery pathsToken limit escalation, context exhaustion auto-compact, model overload fallback, permission deny, tool crash, stuck detection, max turnsDo everything possible to avoid an interruption
Five-layer compressionBudget Reduction → Snip → Microcompact → Context Collapse → Auto-CompactAvoids context explosion

Verification inside the loop:

Anthropic's official architecture document (ArXiv: 2604.14228) states clearly:

"Agents tend to respond by confidently praising the work, even when quality is mediocre, motivating separation of generation from evaluation."

Claude Code embeds verification into three stages of the loop: gather context → take action → verify the result. These are not three separate steps, but an interwoven loop: after a code change, tests run immediately; when a test fails, the error message is read immediately and a fix follows.


2.2 Cursor: Multi-Agent Matrix + Model Race

Loop type: Multi-Agent Orchestration + Model Race Pattern

Five-layer harness architecture:

1. Interface Layer    → Task entry (Slack, GitHub, IDE, Mobile)
2. Orchestration Layer → Plan mode, Model Router, Race Pattern, Subagent Dispatch
3. Execution Layer     → Isolated VM (non-local), multiple agents in parallel
4. Verification Layer  → Each agent can self-verify using Computer Use, Video, Screenshot, Logs
5. Output Layer        → PR + Artifacts (not chat messages)

Model race (Race Pattern):

Cursor's most innovative design is to send the same problem to multiple models at once and select the best result.

Dispatch problem → Model A | Model B | Model C (parallel)
                 → Select best output

The Cursor team reports that this pattern "significantly improves the quality of the final output", especially for difficult bugs that require a small number of precise edits.

Isolated execution for agents:

Cursor automatically creates a separate git worktree for each agent. Each agent edits, builds, and tests code in its own worktree without interfering with the others. This solves the hardest problem for parallel agents: file conflicts.

Model routing:

Composer 2 (Cursor in-house MoE) → routine coding
Claude → reasoning-intensive subtasks
Gemini → large context window processing
GPT-5 Codex → long-running cloud agent

2.3 Aider: Git as Memory, Architect-Editor Separation

Loop type: Single-agent loop with Git-as-memory + Architect/Editor split

Core loop:

User prompt
  → RepoMap (PageRank graph ranking, select the most relevant code context)
  → Assemble context (system prompt + repo map + chat history + files)
  → LLM call
  → Parse edit (whole-file / diff / udiff / search-replace)
  → Apply edit
  → Lint (automatic)
  → Test (optional)
  → If there is an error → Reflection Loop (auto-correct up to max_reflections times)
  → Git auto-commit (one commit per edit)

RepoMap, Aider's signature innovation:

Aider does not stuff the entire codebase into the prompt. Instead:

  1. It builds a symbol graph of the whole repo, where every definition and reference is a node and an edge
  2. It uses PageRank to select the most relevant code context
  3. It dynamically adjusts the amount of context within a token budget

This lets Aider retain an understanding of the global structure in large codebases without exceeding the context window.

Architect + Editor separation:

Phase 1: Architect (strong reasoning model, such as Claude Opus)
  → Analyze the problem → Generate a detailed modification plan → Submit to the user for approval

Phase 2: Editor (fast model, such as Claude Haiku)
  → Execute edit step by step according to the plan → Lint → Test → Auto-correct

Git as memory:

Aider automatically commits every AI edit to git and generates a descriptive commit message. /undo is simply git revert. This makes the entire edit history fully traceable.


2.4 Cline: Native Plan/Act Dual Mode in VS Code

Loop type: Event-driven streaming loop with Plan/Act separation

Core architecture:

VSCode Extension Host
  ├── Controller (State Management Center)
  ├── Task (each task is an independent instance, ensuring isolation)
  │   └── Agent Loop:
  │       while (not done) {
  │         API request → streaming response
  │         → parse XML tool calls
  │         → Safety check (user approval)
  │         → Execute tool
  │         → Git shadow checkpoint
  │         → Update context & state
  │         → Intelligent truncation
  │       }
  └── Webview Provider (React-based UI)

Plan / Act dual mode:

Cline natively supports two separate operating modes:

ModeBehaviorModel Configuration
PlanAnalyze the problem → propose a change plan → user approval. Does not modify any codeStrong reasoning model
ActExecute edits step by step according to the approved plan → test → fixFast model

Technical highlights:

  • XML tool calling: for models that do not support native JSON tool calls, tool call parsing is done using XML format
  • Git shadow versioning: creates an independent git history that does not disturb the user's git history. The agent can freely commit and roll back
  • Generative streaming UI: the tool execution process is visualized in real time, including diffs, browser interactions, and command output
  • Context window intelligence: an intelligent truncation algorithm that preserves semantically critical context

Cline SDK (2026 refactor):

In 2026, Cline extracted its core loop into a standalone @cline/agents package, a stateless runtime loop. This means any application can embed Cline's agent loop.


2.5 SWE-agent: Agent-Computer Interface

Loop type: ReAct-style loop with purpose-built ACI

Core concept: ACI (Agent-Computer Interface)

SWE-agent's core insight is this: an LLM agent is a new kind of end user that needs a purpose-built interface.

Just as humans use an IDE (syntax highlighting, autocomplete, debugger) to assist with programming, an agent needs a purpose-built tool interface.

Loop structure:

for each step:
  Agent generates: Thought + Action
  ACI executes action:
    - find_file / search_file / search_dir (code search, up to 50 results)
    - open / goto (file navigation)
    - edit (including linter feedback: invalid edits are discarded, agent retries)
    - bash (standard Linux commands)
  ACI returns: formatted observation + linter errors
  Agent updates context → next step

Key ACI design decisions:

DesignReason
Search results limited to 50 entriesPrevents the agent from being overwhelmed by too many results
Real-time linter feedback on editSyntax errors are caught immediately, and the agent must fix them before continuing
Invalid edits are discardedPrevents corrupting the codebase
A structured error message formatLets the agent understand the error instead of being confused by it

Performance: SWE-bench pass@1 12.5% (far above all non-interactive LMs at the time), HumanEvalFix 87.7%.


2.6 OpenHands: Built-In Loop Recovery + Security Analyzer

Loop type: State Machine with built-in recovery

Core architecture:

AgentController
  └── step() loop:
      1. Pending Actions? → execute and return
      2. Condensation? → compress history → return condensed view
      3. LLM Query → generate next actions
      4. Safety Check → Security Analyzer (Low/Medium/High)
      5. Confirmation Check → if needs approval, WAITING state
      6. Execute Tools → return observations
      7. StuckDetector → if stuck, attempt_loop_recovery()

StuckDetector + Loop Recovery:

def is_stuck(self) -> bool:
    if self.delegate and self.delegate._is_stuck():
        return True
    return self._stuck_detector.is_stuck(self.headless_mode)

def attempt_loop_recovery(self) -> bool:
    recovery_point = self._stuck_detector.stuck_analysis.loop_start_idx
    # Truncate memory to the loop start point
    await self._truncate_memory_to_point(recovery_point)
    # Restart the agent with the last user message
    await self._restart_with_last_user_message(stuck_analysis)

This is one of very few approaches that treat loop detection and recovery as an architectural mechanism rather than a prompt suggestion.

Security Analyzer with three risk tiers:

RiskBehavior
LowExecute directly
MediumLog a warning + monitor execution
HighBlock execution + request user confirmation

2.7 OpenClaw Loop Engineering: Our In-House System

Loop type: Cron-driven hierarchical multi-agent loop

Three-layer architecture:

Main Agent (Interface)
  └── Receive boss instructions → assign tasks → forward results

Looper Agent (Executor)
  └── cron trigger (every 4h/2h/1h/30min)
      → read LOOP.md + STATE.json
      → execute seven-stage inner loop (max 30 retries)
      → spawn coder-deepseek as checker
      → write to STATE.json

Coder Agents (Workers)
  └── coder-deepseek / coder-qwen / coder-minimax / coder-glm
      → execute specific check / repair / verification tasks

Seven-stage inner loop:

for attempt in 1..30:
  Observe  → Read LOOP.md (immutable checklist) + STATE.json (last state)
  Plan     → update_plan structured steps
  Execute  → 5 mandatory checks + exec/write actual operations
  Verify   → spawn coder-deepseek (independent checker, not self-verify)
  Diagnose → analyze failure cause (timeout / service_down / connection_refused / data_missing)
  Adjust   → escalate according to strategy (1-5 minor / 6-15 medium / 16-30 heavy)
  Retry    → return to Observe, until all PASS or max_retries exhausted

10 auto-healing loops:

LoopScheduleFunction
system-healthEvery 4hFive-item inspection of Qdrant / CC / Cron / OMLX / Tunnel
daily-triageDaily 9AMFollow-ups + email + session triage
cc-babysitterEvery 15minCC session monitoring + escalation
agent-orchestratorEvery 30mincoder-* health checks + load distribution
email-watchEvery 2hEmail scanning + classification + urgent notifications
pr-reviewEvery 4hGitHub PR monitoring + automated review
rnd-discoveryDaily 2AMCodebase TODO/FIXME + pitfall pattern scanning
report-generatorDaily 23:00Comprehensive daily report generation
meaning-discoveryWeeklyMeaning discovery + direction calibration
evolution1st of each monthSelf-evolution + cleanup of outdated content

Methodology enforcement mechanisms:

  • Immutable checklist: the ⛔ block at the top of LOOP.md forbids any agent from adding, removing, or modifying checklist items
  • Gate 0 integrity validation: the checker forcibly verifies that all items have been executed
  • OMLX L3 external verification: score_enforcer.py independently scans STATE.json, and repeated unhealthy status → external alert
  • compliance_score: each STATE.json records the completion level of the seven stages (0-7), and score_enforcer scans hourly

Field lesson: In the production environment, our agent skipped the verify stage 5 times in a row (20 hours) under a purely prompt-based verification mechanism. After the refactor to architectural enforcement, compliance improved significantly. This directly validates the core thesis of this article.


3. Side-by-Side Comparison Matrix

SystemLoop ShapeExecution-Verification SeparationSelf-HealingContext ManagementMulti-ModelComplexity
Claude CodeSingle while(true)Stop Hook + Subagent✅ 7 recovery pathsFive-layer compression❌ SingleMedium
CursorMultiple agents in parallelRace Pattern (comparison of multi-model outputs)❌Independent context per agent✅ Multi-model routerHighest
AiderGit-based loopArchitect+Editor separation✅ Lint+Test reflectionRepoMap (PageRank)✅ Dual-model splitLow
ClineEvent-driven streamingPlan/Act dual mode❌ Requires user interventionIntelligent truncation✅ 33+ providersMedium
SWE-agentReAct-styleACI linter feedback❌ Single-step correctionHistory processors❌ SingleLow
OpenHandsState MachineSecurity Analyzer✅ Loop RecoveryCondenser auto-compress✅ Optional routerHigh
OpenClaw LoopCron-driven hierarchyGate 0 + coder-deepseek checker✅ 30 retries + 4-strategy escalationSTATE.json + iteration_history✅ 5 agentsMedium

4. Deep Analysis of Five Architectural Dimensions

4.1 Loop Shape: One Giant Loop vs. Multi-Stage Separation

StrategyRepresentativeProsCons
Single giant loopClaude Code, ClineSimple, one code path handles all scenarios, improvements automatically benefit everythingHard to separate concerns, context bloats quickly
Architect-editor separationAider, Cline Plan/ActHigher plan quality, faster execution, models can be swappedOne extra API call, the plan may be inaccurate
Multiple agents in parallelCursor, OpenClawDifferent agents for different tasks, isolated execution, parallelizableComplex coordination, difficult state synchronization
Cron-driven hierarchyOpenClawScheduled automation, does not depend on user triggers, independent verification layerHigher latency, not real time

4.2 Execution-Verification Separation: Who Decides the Agent Got It Right?

This is the core thesis of this article and the key point where all the systems diverge.

SystemVerification MechanismEnforcement
Claude CodeStop Hook, an external script that forcibly checks at the turn boundary🔴 Architecturally enforced (blocks the turn)
CursorRace Pattern, multiple models output in parallel and the best is selected🟠 Statistical (non-deterministic)
AiderLint + Test reflection loop, driven by hard signals🟠 Configurable
ClineManual user approval (Plan/Act)🟡 Human-driven
SWE-agentACI linter, invalid edits are discarded automatically🟠 Tool layer
OpenHandsSecurity Analyzer, three-tier risk assessment🟠 Configurable
OpenClawGate 0 + coder-deepseek + L3 score_enforcer🔴 Three-layer enforcement

Core insight: The enforcement of verification forms a continuous spectrum from "human-driven" (the user checks it themselves) to "architecturally enforced" (Stop Hook, Gate 0). The closer to the architecturally enforced end, the more reliable the system, but the higher the implementation cost.

4.3 Self-Healing: Can the Agent Recover from Its Own Errors?

SystemRecovery MechanismDegree of Autonomy
Claude Code7 recovery paths (classified handling of various errors such as token/context/model/tool)🔴 Fully automatic
OpenHandsStuckDetector → truncate memory back to the loop start point → restart🔴 Fully automatic
AiderLint error → reflection loop → automatic correction (up to max_reflections times)🟠 Semi-automatic
OpenClaw30 retries + four-level escalating strategy (mild → moderate → heavy adjustment)🔴 Fully automatic

4.4 Context Management: How to Prevent the Context Window from Exploding?

SystemStrategy
Claude CodeFive-layer compression pipeline: Budget → Snip → Microcompact → Collapse → Auto-Compact
CursorIndependent context per agent + only inject relevant code snippets (semantic search)
AiderRepoMap (PageRank graph ranking, sends only the most relevant symbol definitions)
ClineIntelligent truncation, keeps semantically critical parts and discards duplicated content
SWE-agentHistory processors + structured error messages (suppresses redundant output)
OpenClawSTATE.json persistence + iteration_history, an independent session per run, not relying on accumulated context

4.5 Multi-Model Strategy: One Model to Rule Them All vs. a Diverse Ecosystem

StrategySystemApplicable Scenario
Single strongest modelClaude Code, SWE-agentA single task type, and the model's capability is sufficient
Router (dynamic allocation)Cursor Auto modeDifferent subtasks require different model characteristics
Race (parallel competition)Cursor Race PatternDifficult bugs requiring a small number of precise edits
Architect+Editor (split into phases)Aider, Cline Plan/ActComplex multi-step tasks
Dedicated agent poolOpenClaw (5 agents)Different agents handle different domains

5. From Comparison to Practice: Our Choices and Lessons

5.1 Why We Chose a Cron-Driven Hierarchy

Our requirements differ from mainstream IDE agents:

RequirementIDE Agent (Claude Code/Cursor)Our System (OpenClaw Loop)
TriggerUser enters a promptScheduled automation (cron)
Time scaleSeconds to minutesHours to days
Task typeWriting code, fixing bugsSystem inspection, email monitoring, agent coordination
Verification requirementTests passService health + compliance audit
Reliability requirementOccasional failure is acceptableContinuous monitoring must not be interrupted

The traditional "user triggers → agent executes → user verifies" loop cannot meet our needs. We need an automated system that does not depend on a human in the loop.

5.2 The Most Important Lessons We Learned from This Comparison

Lesson one: Claude Code's Stop Hook is the simplest and most powerful verification mechanism. We have ported its idea into Gate 0 and the L3 score_enforcer.

Lesson two: Cursor's Race Pattern hints at the future direction. For high-risk decisions (such as determining OMLX service status), having multiple models judge at the same time and then comparing the results may be more reliable than a single checker.

Lesson three: Aider's RepoMap is a model of context management. Instead of stuffing everything into the prompt, it intelligently selects the most relevant context. Our LOOP.md + STATE.json pattern was inspired by exactly this.

Lesson four: OpenHands' Loop Recovery is what we plan to implement next. Upgrade stuck detection from a cron timeout into genuine behavioral pattern detection.

Lesson five: SWE-agent's ACI philosophy reminds us that tool design is interface design. Our LOOP.md format, STATE.json schema, and cron prompt structure are all essentially an Agent-Computer Interface.


6. Conclusion

In 2026, AI Agent systems are moving from the "demo stage" into the "production stage". In this transition, the choice of loop architecture matters more than the choice of model.

The comparison of seven systems reveals several design principles that recur again and again:

  1. The loop shape must match the task characteristics. Short-term coding tasks suit a single while(true); long-term monitoring tasks suit a cron-driven hierarchy.

  2. Verification cannot rely on good intentions. Claude Code's Stop Hook, Cursor's Race Pattern, and our Gate 0 are all architecturally enforced, not prompt suggestions.

  3. Context management determines the ceiling. Aider's RepoMap, Claude Code's five-layer compression, and Cursor's isolated context: in large-scale tasks, context management capability directly determines how far an agent can go.

  4. Multi-model strategies are a competitive advantage. Whether it is Cursor's Race Pattern, Aider's Architect+Editor, or our dedicated agent pool, combining multiple models is almost always better than a single model.

  5. Self-healing must be an architectural feature. OpenHands' Loop Recovery and Claude Code's 7 recovery paths point in the right direction: write the recovery logic as code, not as prompt text.


This article is based on a comprehensive Loop Engineering audit conducted in the Junze Zhiku production environment on June 14, 2026. All code references come from the corresponding projects' public source code or official documentation.