Agentic Research

Comparison of AI Agent Verification Architectures: From Prompt Self-Awareness to Architectural Enforcement, a Source-Level Analysis of Seven Approaches

2026/06/1461 min readBryan Chan閱讀中文原文
TopicsAI AgentLoop EngineeringVerificationArchitectureClaude Code

Core proposition: Once an AI Agent evolves from "single-shot Q&A" into "autonomous multi-step execution", a fatal question emerges: who verifies that every step the Agent takes is correct? "Self-checking" at the pure prompt level is inherently unstable and must be enforced by the architecture layer. Approaches covered: Claude Code · MetaGPT · OpenHands · VIGIL · PEV · ReVeal · odot Methodology: Official documentation analysis + source-level dissection + production environment field verification


Abstract

Once an AI Agent evolves from "single-shot Q&A" into "autonomous multi-step execution", a fatal question emerges: who verifies that every step the Agent takes is correct? This article compares seven mainstream Agent verification architecture approaches on the market, Claude Code, MetaGPT, OpenHands, VIGIL, Planner-Executor-Verifier, ReVeal, and odot, performing a technical analysis along four dimensions: execution-verification separation, enforcement, self-healing capability, and external supervision. The study finds that the most reliable approaches all follow the same iron rule: the executing agent and the verifying agent must be different entities. "Self-checking" at the pure prompt level is inherently unstable and must be enforced by the architecture layer.


1. Problem Definition: Why Agents Need Independent Verification

In multi-step Agent execution, every action changes the system state. A tiny error that occurs at step 3 may not surface until step 12. More dangerously: after finishing a step, an Agent tends to "declare success" rather than "verify success".

We observed a typical case in the production environment: across 5 consecutive health inspections, DeepSeek V4 Pro detected that a critical service was down (OMLX, the local MLX model inference service, 7 models, ~169GB), and the configuration file explicitly marked that service as 🔴 high severity, yet the model still independently judged it to be a "non-critical service that needs no action", leaving a known fault unrepaired for 20 consecutive hours.

This is not a problem of a single model. It exposes a systemic limitation of all current LLM architectures: models lack metacognition, so when internal reasoning conflicts with external rules, the model may independently override the rules.

Independent verification architectures were created precisely to solve this problem.


2. Technical Comparison of Seven Approaches

2.1 Claude Code (Anthropic): Minimal Design, Maximal Constraint

Architecture type: Single Agent loop + external enforced verification

Core mechanisms:

MechanismImplementationEnforcement
Stop HookBefore every turn ends, an external script checks the result. If it does not pass, the turn is blocked, up to 8 blocks🔴 Architecturally enforced
Verification SubagentAn independent subagent reviews the main agent's output using a fresh model call (not contaminated by the main agent's context)🟠 User-optional
/goal ConditionSets a termination condition, and an independent evaluator re-evaluates after every turn🟠 User-optional
Dynamic WorkflowsJS scripts orchestrate multiple subagents for adversarial cross-review🟡 Advanced feature

Technical highlights:

Claude Code's design philosophy is "minimal abstraction, maximal enforcement". Its Stop Hook mechanism is extremely simple, just an external script that returns pass/fail, but its effect is extremely strong: the agent cannot skip verification, because the turn simply will not end.

Hooks run inside the application process and do not consume the agent context window. This is a key design decision: verification logic is isolated from the reasoning context.

Anthropic's official documentation states clearly:

"A verification subagent that checks findings has a fresh model, so the agent doing the work isn't also the judge."

In translation: the executor cannot also be the judge.

Source-level implementation:

Claude Code's core is an AsyncGenerator loop. The flow of each turn is:

LLM reasoning → generate tool_call → execute tool → collect results
  → check token budget → check Stop Hook → continue or end

Compared with the design of most frameworks, Claude Code's strategy is the most "brute-force": it does not suggest self-checking in the prompt, but inserts an unavoidable gate at the turn boundary.


2.2 MetaGPT (ICLR 2024 Oral, 45K⭐): Multi-Role Pipeline

Architecture type: Multi-Agent SOP pipeline, role isolation

Core role chain:

Product Manager → Architect → Project Manager → Engineer → QA Engineer
     Requirements Analysis         System Design         Task Breakdown        Write Code       Write Tests+Verification

Technical mechanisms:

MechanismImplementation
Role isolation5 independent Agents, each with its own prompt, memory, and tools
Publish-SubscribeUpstream output = the downstream's only input. Each role _watches specific upstream Actions, with no other information source
QA EngineerA dedicated verification role, independent of the Engineer, and not allowed to write production code
Executable FeedbackThe Engineer generates unit tests → executes → gets errors → fixes. But final verification is completed independently by the QA Engineer

Enforcement comes from structural isolation:

The QA Engineer cannot access the Engineer's internal reasoning and can only see the code and test results the Engineer outputs. This information asymmetry is precisely the advantage: QA reviews the output with a "fresh pair of eyes".

ICLR paper data:

MetricWithout QA roleWith QA role
HumanEval Pass@1Baseline+4.2%
MBPP Pass@1Baseline+5.4%
Manual modification cost2.250.83 (-63%)

2.3 OpenHands (All Hands AI, 48K⭐): Built-In Loop Recovery

Architecture type: State Machine + Security Analyzer + Loop Recovery

Technical architecture:

ComponentFunction
CodeActAgentA stateless step() loop: each call invokes the LLM → executes tools → returns results
Security AnalyzerEvery action passes a three-tier risk assessment before execution: Low (execute directly) / Medium (log + monitor) / High (block + request confirmation)
StuckDetectorDetects whether the agent is stuck in a repetitive behavior loop
Loop RecoveryOnce a loop is detected: truncate memory back to the loop start point → restart the agent → inject a new repair instruction
CondenserAutomatically compresses history while preserving key decisions

StuckDetector source snippet:

def is_stuck(self) -> bool:
    """Checks if the agent or its delegate is stuck in a loop."""
    if self.delegate and self.delegate._is_stuck():
        return True
    return self._stuck_detector.is_stuck(self.headless_mode)

def attempt_loop_recovery(self) -> bool:
    if not self._stuck_detector.stuck_analysis:
        return False
    recovery_point = self._stuck_detector.stuck_analysis.loop_start_idx
    # Truncate memory to the recovery point
    await self._truncate_memory_to_point(recovery_point)
    # Restart agent with last user message
    await self._restart_with_last_user_message(stuck_analysis)

Design highlights:

OpenHands' distinctive feature is that it elevates loop detection and recovery into an architectural mechanism (Python code). Most Agent frameworks at best "limit the maximum number of steps" to prevent infinite loops, but OpenHands genuinely achieves detect the loop → diagnose the loop start point → truncate memory → start over. The Security Analyzer's three-tier risk model is also worth noting: it strips security checks out of the prompt layer and turns them into an independent gate before execution.


2.4 VIGIL (ArXiv 2025): External Reflective Supervision

Architecture type: Out-of-band Reflective Runtime

Core idea:

"Rather than embedding self-repair logic within an agent's inference cycle, VIGIL persists as an external, introspective layer-detecting latent faults, identifying structural breakdowns, and proposing concrete fixes to prompt and code."

In translation: rather than embedding self-repair logic inside the agent's inference loop, make it an external, independent reflection layer.

Pipeline architecture:

Main Agent (Task Execution)
      ↓ Generates logs
VIGIL Runtime (External Supervision Layer, independent process)
      ↓ Reads behavior logs
Emotional Bank (Emotional Memory Bank, with time decay strategy)
      ↓ Emotional assessment
RBT Diagnosis (Roses/Buds/Thorns structured diagnosis)
      ↓
Prompt Diff + Code Diff (Generates repair suggestions, but does not execute automatically)

Stage-Gate state machine:

start → eb_updated → diagnosed → prompt_done → diff_done

Illegal state transitions raise an error immediately, giving the LLM no opportunity to arrange the tool sequence on its own.

Self-healing capability:

The most instructive moment in VIGIL is this: when its own diagnostic tool failed due to a schema mismatch, VIGIL automatically detected this internal error, generated a fallback diagnosis, and issued a repair plan for itself.

This demonstrates a kind of "meta-procedural self-repair", in which even the supervisory layer itself can be supervised. This is a rare instance in the current literature.


2.5 Planner-Executor-Verifier (Agent Patterns Catalog): Pattern-Level Abstraction

Architecture type: Design pattern (not a specific implementation)

Core constraint:

"No plan step's effect is accepted without an independent verifier check; same-model self-verify is excluded."

This may be the most concise and most uncompromising verification rule in the entire field.

Key differences from Plan-Execute:

Plan-ExecutePlanner-Executor-Verifier
Step verificationNone, or the executor self-checksAn independent Verifier checks every step
Failure detectionDetects tool errors onlyDetects goal-progress drift
Correction mechanismRetries the same stepReplan (re-plan after accounting for drift)
Applicable scenarioSimple linear tasksMulti-step state-changing tasks

Classic case:

Plan: "Refactor module X" (12 steps)
Step 4: rename function → Executor executes → Tool returns success
PEV Verifier checks:
  - "Does the refactor still compile?"
  - "Do all callers still resolve?"
→ Found 3 unresolved callers
→ Replan triggered (call-site list as new constraint)
→ Continue executing the corrected plan

Plan-Execute without PEV:
→ Executor sees step 4 success → goes directly to step 5
→ Only at step 12 discovers the entire refactor has 3 broken callers
→ Needs to roll back 8 steps

PEV's insight is this: tool-level success ≠ goal-level progress. A tool returning success does not mean the plan is advancing.


2.6 ReVeal (ArXiv 2025): RL-Driven Generation-Verification

Architecture type: Multi-turn RL with Execution Feedback

Training pipeline:

StageBehavior
GenerationThe model generates candidate code based on the task description
Test GenerationThe same model generates test cases for its own code
ExecutionExecutes the test cases in a sandbox
VerificationObtains the pass/fail hard signal
RL UpdateUpdates the model weights based on pass/fail (RL-zero, no distillation, no SFT)

What is distinctive:

ReVeal is the only approach that solves the verification problem with RL training rather than prompt engineering. It lets the same model learn both "generating code" and "writing tests for code".

The key difference: the verification signal comes from executable test results (a hard signal), not from the LLM's textual judgment (a soft signal). This fundamentally avoids the problem of "the model rationalizing itself": the code either passes the test or it does not, with no gray area.


2.7 odot (Frobenius Condition): A Mathematically Guaranteed Closed Loop

Architecture type: Mathematical verification closed loop (Frobenius Condition)

Core mechanism:

THINK → ACT → OBSERVE (including mandatory verification step μ) → UPDATE

Frobenius Condition: μ(δ(query)) == query must hold
→ Holds (Frobenius CLOSED): loop continues
→ Does not hold (Frobenius OPEN): the model receives an explicit failure message, must correct and retry, and must not skip it

Design features:

odot inserts a mandatory verification step immediately after every tool call. This verification is not optional: if the Frobenius Condition does not hold (μ(δ(q)) != q), the loop cannot advance.

This is spiritually consistent with Claude Code's Stop Hook: replacing prompt-level suggestions with mathematical and code-level enforcement.

odot also supports B4 Belnap FOUR four-valued logic verification: N (Neither), T (True), F (False), B (Both/contradiction). In paraconsistent logic, B does not cause an explosion, which allows odot to continue gracefully when facing contradictory information.


3. Side-by-Side Comparison Matrix

ApproachExecution-Verification SeparationEnforcement MechanismSelf-HealingExternal SupervisionVerification SignalComplexity
Claude Code✅ Subagent Hook🔴 Architecturally enforced (block turn)✅ Iterative correction✅ Hook+Subagentpass/fail scriptLow
MetaGPT✅ QA Engineer🔴 Role isolation (publish-subscribe)✅ Executable Feedback❌ Internal roleTest resultsMedium
OpenHands❌ Same Agent🟠 Security Analyzer (configurable)✅ Loop Recovery❌ Internal mechanismTool return valueHigh
VIGIL✅ External reflection layer🔴 Stage-Gate (illegal transitions raise errors)✅ Self-diagnosis + repair✅ Independent processBehavioral log analysisHigh
PEV✅ Independent Verifier🔴 Pattern-enforced (no verification, no acceptance)✅ Replan on failure✅ External Verifiergoal-progressMedium
ReVeal❌ Same model🟡 RL training✅ Multi-turn iteration❌ Internal loopTest hard signalHighest
odot❌ Same model🔴 Frobenius Condition✅ Closed-loop retry❌ InternalMathematical equationLow
Our Loop v1❌ 0 spawn🟢 Prompt text❌ Zero action❌ NoneNoneLow

4. Design Patterns Worth Borrowing

4.1 The Stop Hook Pattern (borrowed from Claude Code)

After each loop finishes executing and before writing the final result, forcibly run an external check:

def stop_hook(loop_id, result):
    if not result.methodology_compliance.verify.done:
        return False, "VERIFY phase not executed"
    if result.methodology_compliance.verify.checker_session is None:
        return False, "No independent checker spawned"
    for alert in result.alerts:
        if alert.severity == "critical" and not alert.escalated:
            return False, f"Critical alert not escalated: {alert.message}"
    return True, ""

4.2 The Role Isolation Pattern (borrowed from MetaGPT)

Split each loop into two independent cron jobs:

Job A: loop-xxx-execute (looper agent)
  → Observe → Plan → Execute → Write to draft STATE.json

Job B: loop-xxx-verify (coder-deepseek agent)
  → Read draft STATE.json → Independently Verify → If passed, merge
  → If not passed → Write fix_instructions → Trigger Job A retry

Key constraint: the looper never writes the final STATE.json; only after coder-deepseek's verify passes can it be written.

4.3 The External Reflection Layer Pattern (borrowed from VIGIL)

Extend the existing score_enforcer.py into a complete external supervision layer:

class LoopSupervisor:
    def scan_all_loops(self):
        for loop in get_active_loops():
            self.check_methodology_compliance(loop)
            self.check_omlx_l3(loop)
            self.check_escalation_compliance(loop)

    def check_methodology_compliance(self, loop):
        state = load_state(loop)
        score = state.methodology_compliance.compliance_score
        if score < 4:
            alert(f"Loop {loop} methodology score {score}/7: possible skip")

5. Conclusion

5.1 Core Findings

This article analyzed seven AI Agent verification architecture approaches and found that they share a recurring design principle:

"The executor cannot also be the judge."

  • Claude Code uses the Stop Hook to ensure an external check: not a prompt suggestion, but a hard gate at the turn boundary
  • MetaGPT uses a QA Engineer to achieve role isolation: the Engineer cannot verify its own code, and the QA cannot write production code
  • VIGIL uses an external reflection layer to monitor the main Agent, and even the supervisory layer's own failures can be self-diagnosed
  • PEV explicitly forbids same-model self-verify; a step without an independent verifier is not accepted
  • odot uses the Frobenius Condition to enforce a closed loop at the mathematical level; if it does not hold, the loop is not allowed to advance
  • ReVeal, although it uses the same model, takes its verification signal from executable tests rather than textual judgment

Pure prompt-level "self-checking", such as writing "please verify your own result" in the prompt, proved unstable in every comparison.

5.2 Field Verification

We verified this conclusion in our production Loop Engineering system: when verification existed only in prompt text, the agent skipped 100% of verify stages across 5 consecutive runs (0/5 spawned checker). But once we embedded the verify mechanism into the cron structure and the forced STATE.json format, the compliance rate improved significantly.

5.3 Four Recommendations

For teams building production-grade Agent systems:

  1. Never let the same agent verify itself. Use a different agent, a different model call, or an external hook. Claude Code and MetaGPT have already proven the necessity of this rule.

  2. The verification signal should be a hard signal. Executable tests, mathematical equations (Frobenius Condition), or structured schema checks are superior to LLM textual judgment. A hard signal has no gray area.

  3. Enforcement must come from the architecture, not the prompt. Stop Hook, Stage Gate, Role Isolation, and Frobenius Condition are all mechanisms that an LLM cannot "rationalize its way past". Prompt text can always be "reinterpreted".

  4. An external supervision layer is the best long-term solution. VIGIL's out-of-band reflective runtime pattern is currently the most scalable approach: it does not intrude on existing agent logic and does not consume the context window, yet it provides independent, continuous compliance monitoring. Our score_enforcer.py is evolving in this direction.


This article is based on a comprehensive Loop Engineering audit conducted in the Junze Zhiku production environment on June 14, 2026. All code references come from the corresponding projects' public source code or official documentation. Special thanks to the DeepSeek team for providing an excellent model; it was precisely its "failures" that allowed us to dig deeper into the underlying problems of this field.