Comparison of AI Agent Verification Architectures: From Prompt Self-Awareness to Architectural Enforcement, a Source-Level Analysis of Seven Approaches
Core proposition: Once an AI Agent evolves from "single-shot Q&A" into "autonomous multi-step execution", a fatal question emerges: who verifies that every step the Agent takes is correct? "Self-checking" at the pure prompt level is inherently unstable and must be enforced by the architecture layer. Approaches covered: Claude Code · MetaGPT · OpenHands · VIGIL · PEV · ReVeal · odot Methodology: Official documentation analysis + source-level dissection + production environment field verification
Abstract
Once an AI Agent evolves from "single-shot Q&A" into "autonomous multi-step execution", a fatal question emerges: who verifies that every step the Agent takes is correct? This article compares seven mainstream Agent verification architecture approaches on the market, Claude Code, MetaGPT, OpenHands, VIGIL, Planner-Executor-Verifier, ReVeal, and odot, performing a technical analysis along four dimensions: execution-verification separation, enforcement, self-healing capability, and external supervision. The study finds that the most reliable approaches all follow the same iron rule: the executing agent and the verifying agent must be different entities. "Self-checking" at the pure prompt level is inherently unstable and must be enforced by the architecture layer.
1. Problem Definition: Why Agents Need Independent Verification
In multi-step Agent execution, every action changes the system state. A tiny error that occurs at step 3 may not surface until step 12. More dangerously: after finishing a step, an Agent tends to "declare success" rather than "verify success".
We observed a typical case in the production environment: across 5 consecutive health inspections, DeepSeek V4 Pro detected that a critical service was down (OMLX, the local MLX model inference service, 7 models, ~169GB), and the configuration file explicitly marked that service as 🔴 high severity, yet the model still independently judged it to be a "non-critical service that needs no action", leaving a known fault unrepaired for 20 consecutive hours.
This is not a problem of a single model. It exposes a systemic limitation of all current LLM architectures: models lack metacognition, so when internal reasoning conflicts with external rules, the model may independently override the rules.
Independent verification architectures were created precisely to solve this problem.
2. Technical Comparison of Seven Approaches
2.1 Claude Code (Anthropic): Minimal Design, Maximal Constraint
Architecture type: Single Agent loop + external enforced verification
Core mechanisms:
| Mechanism | Implementation | Enforcement |
|---|---|---|
| Stop Hook | Before every turn ends, an external script checks the result. If it does not pass, the turn is blocked, up to 8 blocks | 🔴 Architecturally enforced |
| Verification Subagent | An independent subagent reviews the main agent's output using a fresh model call (not contaminated by the main agent's context) | 🟠 User-optional |
/goal Condition | Sets a termination condition, and an independent evaluator re-evaluates after every turn | 🟠 User-optional |
| Dynamic Workflows | JS scripts orchestrate multiple subagents for adversarial cross-review | 🟡 Advanced feature |
Technical highlights:
Claude Code's design philosophy is "minimal abstraction, maximal enforcement". Its Stop Hook mechanism is extremely simple, just an external script that returns pass/fail, but its effect is extremely strong: the agent cannot skip verification, because the turn simply will not end.
Hooks run inside the application process and do not consume the agent context window. This is a key design decision: verification logic is isolated from the reasoning context.
Anthropic's official documentation states clearly:
"A verification subagent that checks findings has a fresh model, so the agent doing the work isn't also the judge."
In translation: the executor cannot also be the judge.
Source-level implementation:
Claude Code's core is an AsyncGenerator loop. The flow of each turn is:
LLM reasoning → generate tool_call → execute tool → collect results
→ check token budget → check Stop Hook → continue or end
Compared with the design of most frameworks, Claude Code's strategy is the most "brute-force": it does not suggest self-checking in the prompt, but inserts an unavoidable gate at the turn boundary.
2.2 MetaGPT (ICLR 2024 Oral, 45K⭐): Multi-Role Pipeline
Architecture type: Multi-Agent SOP pipeline, role isolation
Core role chain:
Product Manager → Architect → Project Manager → Engineer → QA Engineer
Requirements Analysis System Design Task Breakdown Write Code Write Tests+Verification
Technical mechanisms:
| Mechanism | Implementation |
|---|---|
| Role isolation | 5 independent Agents, each with its own prompt, memory, and tools |
| Publish-Subscribe | Upstream output = the downstream's only input. Each role _watches specific upstream Actions, with no other information source |
| QA Engineer | A dedicated verification role, independent of the Engineer, and not allowed to write production code |
| Executable Feedback | The Engineer generates unit tests → executes → gets errors → fixes. But final verification is completed independently by the QA Engineer |
Enforcement comes from structural isolation:
The QA Engineer cannot access the Engineer's internal reasoning and can only see the code and test results the Engineer outputs. This information asymmetry is precisely the advantage: QA reviews the output with a "fresh pair of eyes".
ICLR paper data:
| Metric | Without QA role | With QA role |
|---|---|---|
| HumanEval Pass@1 | Baseline | +4.2% |
| MBPP Pass@1 | Baseline | +5.4% |
| Manual modification cost | 2.25 | 0.83 (-63%) |
2.3 OpenHands (All Hands AI, 48K⭐): Built-In Loop Recovery
Architecture type: State Machine + Security Analyzer + Loop Recovery
Technical architecture:
| Component | Function |
|---|---|
| CodeActAgent | A stateless step() loop: each call invokes the LLM → executes tools → returns results |
| Security Analyzer | Every action passes a three-tier risk assessment before execution: Low (execute directly) / Medium (log + monitor) / High (block + request confirmation) |
| StuckDetector | Detects whether the agent is stuck in a repetitive behavior loop |
| Loop Recovery | Once a loop is detected: truncate memory back to the loop start point → restart the agent → inject a new repair instruction |
| Condenser | Automatically compresses history while preserving key decisions |
StuckDetector source snippet:
def is_stuck(self) -> bool:
"""Checks if the agent or its delegate is stuck in a loop."""
if self.delegate and self.delegate._is_stuck():
return True
return self._stuck_detector.is_stuck(self.headless_mode)
def attempt_loop_recovery(self) -> bool:
if not self._stuck_detector.stuck_analysis:
return False
recovery_point = self._stuck_detector.stuck_analysis.loop_start_idx
# Truncate memory to the recovery point
await self._truncate_memory_to_point(recovery_point)
# Restart agent with last user message
await self._restart_with_last_user_message(stuck_analysis)
Design highlights:
OpenHands' distinctive feature is that it elevates loop detection and recovery into an architectural mechanism (Python code). Most Agent frameworks at best "limit the maximum number of steps" to prevent infinite loops, but OpenHands genuinely achieves detect the loop → diagnose the loop start point → truncate memory → start over. The Security Analyzer's three-tier risk model is also worth noting: it strips security checks out of the prompt layer and turns them into an independent gate before execution.
2.4 VIGIL (ArXiv 2025): External Reflective Supervision
Architecture type: Out-of-band Reflective Runtime
Core idea:
"Rather than embedding self-repair logic within an agent's inference cycle, VIGIL persists as an external, introspective layer-detecting latent faults, identifying structural breakdowns, and proposing concrete fixes to prompt and code."
In translation: rather than embedding self-repair logic inside the agent's inference loop, make it an external, independent reflection layer.
Pipeline architecture:
Main Agent (Task Execution)
↓ Generates logs
VIGIL Runtime (External Supervision Layer, independent process)
↓ Reads behavior logs
Emotional Bank (Emotional Memory Bank, with time decay strategy)
↓ Emotional assessment
RBT Diagnosis (Roses/Buds/Thorns structured diagnosis)
↓
Prompt Diff + Code Diff (Generates repair suggestions, but does not execute automatically)
Stage-Gate state machine:
start → eb_updated → diagnosed → prompt_done → diff_done
Illegal state transitions raise an error immediately, giving the LLM no opportunity to arrange the tool sequence on its own.
Self-healing capability:
The most instructive moment in VIGIL is this: when its own diagnostic tool failed due to a schema mismatch, VIGIL automatically detected this internal error, generated a fallback diagnosis, and issued a repair plan for itself.
This demonstrates a kind of "meta-procedural self-repair", in which even the supervisory layer itself can be supervised. This is a rare instance in the current literature.
2.5 Planner-Executor-Verifier (Agent Patterns Catalog): Pattern-Level Abstraction
Architecture type: Design pattern (not a specific implementation)
Core constraint:
"No plan step's effect is accepted without an independent verifier check; same-model self-verify is excluded."
This may be the most concise and most uncompromising verification rule in the entire field.
Key differences from Plan-Execute:
| Plan-Execute | Planner-Executor-Verifier | |
|---|---|---|
| Step verification | None, or the executor self-checks | An independent Verifier checks every step |
| Failure detection | Detects tool errors only | Detects goal-progress drift |
| Correction mechanism | Retries the same step | Replan (re-plan after accounting for drift) |
| Applicable scenario | Simple linear tasks | Multi-step state-changing tasks |
Classic case:
Plan: "Refactor module X" (12 steps)
Step 4: rename function → Executor executes → Tool returns success
PEV Verifier checks:
- "Does the refactor still compile?"
- "Do all callers still resolve?"
→ Found 3 unresolved callers
→ Replan triggered (call-site list as new constraint)
→ Continue executing the corrected plan
Plan-Execute without PEV:
→ Executor sees step 4 success → goes directly to step 5
→ Only at step 12 discovers the entire refactor has 3 broken callers
→ Needs to roll back 8 steps
PEV's insight is this: tool-level success ≠ goal-level progress. A tool returning success does not mean the plan is advancing.
2.6 ReVeal (ArXiv 2025): RL-Driven Generation-Verification
Architecture type: Multi-turn RL with Execution Feedback
Training pipeline:
| Stage | Behavior |
|---|---|
| Generation | The model generates candidate code based on the task description |
| Test Generation | The same model generates test cases for its own code |
| Execution | Executes the test cases in a sandbox |
| Verification | Obtains the pass/fail hard signal |
| RL Update | Updates the model weights based on pass/fail (RL-zero, no distillation, no SFT) |
What is distinctive:
ReVeal is the only approach that solves the verification problem with RL training rather than prompt engineering. It lets the same model learn both "generating code" and "writing tests for code".
The key difference: the verification signal comes from executable test results (a hard signal), not from the LLM's textual judgment (a soft signal). This fundamentally avoids the problem of "the model rationalizing itself": the code either passes the test or it does not, with no gray area.
2.7 odot (Frobenius Condition): A Mathematically Guaranteed Closed Loop
Architecture type: Mathematical verification closed loop (Frobenius Condition)
Core mechanism:
THINK → ACT → OBSERVE (including mandatory verification step μ) → UPDATE
Frobenius Condition: μ(δ(query)) == query must hold
→ Holds (Frobenius CLOSED): loop continues
→ Does not hold (Frobenius OPEN): the model receives an explicit failure message, must correct and retry, and must not skip it
Design features:
odot inserts a mandatory verification step immediately after every tool call. This verification is not optional: if the Frobenius Condition does not hold (μ(δ(q)) != q), the loop cannot advance.
This is spiritually consistent with Claude Code's Stop Hook: replacing prompt-level suggestions with mathematical and code-level enforcement.
odot also supports B4 Belnap FOUR four-valued logic verification: N (Neither), T (True), F (False), B (Both/contradiction). In paraconsistent logic, B does not cause an explosion, which allows odot to continue gracefully when facing contradictory information.
3. Side-by-Side Comparison Matrix
| Approach | Execution-Verification Separation | Enforcement Mechanism | Self-Healing | External Supervision | Verification Signal | Complexity |
|---|---|---|---|---|---|---|
| Claude Code | ✅ Subagent Hook | 🔴 Architecturally enforced (block turn) | ✅ Iterative correction | ✅ Hook+Subagent | pass/fail script | Low |
| MetaGPT | ✅ QA Engineer | 🔴 Role isolation (publish-subscribe) | ✅ Executable Feedback | ❌ Internal role | Test results | Medium |
| OpenHands | ❌ Same Agent | 🟠 Security Analyzer (configurable) | ✅ Loop Recovery | ❌ Internal mechanism | Tool return value | High |
| VIGIL | ✅ External reflection layer | 🔴 Stage-Gate (illegal transitions raise errors) | ✅ Self-diagnosis + repair | ✅ Independent process | Behavioral log analysis | High |
| PEV | ✅ Independent Verifier | 🔴 Pattern-enforced (no verification, no acceptance) | ✅ Replan on failure | ✅ External Verifier | goal-progress | Medium |
| ReVeal | ❌ Same model | 🟡 RL training | ✅ Multi-turn iteration | ❌ Internal loop | Test hard signal | Highest |
| odot | ❌ Same model | 🔴 Frobenius Condition | ✅ Closed-loop retry | ❌ Internal | Mathematical equation | Low |
| Our Loop v1 | ❌ 0 spawn | 🟢 Prompt text | ❌ Zero action | ❌ None | None | Low |
4. Design Patterns Worth Borrowing
4.1 The Stop Hook Pattern (borrowed from Claude Code)
After each loop finishes executing and before writing the final result, forcibly run an external check:
def stop_hook(loop_id, result):
if not result.methodology_compliance.verify.done:
return False, "VERIFY phase not executed"
if result.methodology_compliance.verify.checker_session is None:
return False, "No independent checker spawned"
for alert in result.alerts:
if alert.severity == "critical" and not alert.escalated:
return False, f"Critical alert not escalated: {alert.message}"
return True, ""
4.2 The Role Isolation Pattern (borrowed from MetaGPT)
Split each loop into two independent cron jobs:
Job A: loop-xxx-execute (looper agent)
→ Observe → Plan → Execute → Write to draft STATE.json
Job B: loop-xxx-verify (coder-deepseek agent)
→ Read draft STATE.json → Independently Verify → If passed, merge
→ If not passed → Write fix_instructions → Trigger Job A retry
Key constraint: the looper never writes the final STATE.json; only after coder-deepseek's verify passes can it be written.
4.3 The External Reflection Layer Pattern (borrowed from VIGIL)
Extend the existing score_enforcer.py into a complete external supervision layer:
class LoopSupervisor:
def scan_all_loops(self):
for loop in get_active_loops():
self.check_methodology_compliance(loop)
self.check_omlx_l3(loop)
self.check_escalation_compliance(loop)
def check_methodology_compliance(self, loop):
state = load_state(loop)
score = state.methodology_compliance.compliance_score
if score < 4:
alert(f"Loop {loop} methodology score {score}/7: possible skip")
5. Conclusion
5.1 Core Findings
This article analyzed seven AI Agent verification architecture approaches and found that they share a recurring design principle:
"The executor cannot also be the judge."
- Claude Code uses the Stop Hook to ensure an external check: not a prompt suggestion, but a hard gate at the turn boundary
- MetaGPT uses a QA Engineer to achieve role isolation: the Engineer cannot verify its own code, and the QA cannot write production code
- VIGIL uses an external reflection layer to monitor the main Agent, and even the supervisory layer's own failures can be self-diagnosed
- PEV explicitly forbids same-model self-verify; a step without an independent verifier is not accepted
- odot uses the Frobenius Condition to enforce a closed loop at the mathematical level; if it does not hold, the loop is not allowed to advance
- ReVeal, although it uses the same model, takes its verification signal from executable tests rather than textual judgment
Pure prompt-level "self-checking", such as writing "please verify your own result" in the prompt, proved unstable in every comparison.
5.2 Field Verification
We verified this conclusion in our production Loop Engineering system: when verification existed only in prompt text, the agent skipped 100% of verify stages across 5 consecutive runs (0/5 spawned checker). But once we embedded the verify mechanism into the cron structure and the forced STATE.json format, the compliance rate improved significantly.
5.3 Four Recommendations
For teams building production-grade Agent systems:
-
Never let the same agent verify itself. Use a different agent, a different model call, or an external hook. Claude Code and MetaGPT have already proven the necessity of this rule.
-
The verification signal should be a hard signal. Executable tests, mathematical equations (Frobenius Condition), or structured schema checks are superior to LLM textual judgment. A hard signal has no gray area.
-
Enforcement must come from the architecture, not the prompt. Stop Hook, Stage Gate, Role Isolation, and Frobenius Condition are all mechanisms that an LLM cannot "rationalize its way past". Prompt text can always be "reinterpreted".
-
An external supervision layer is the best long-term solution. VIGIL's out-of-band reflective runtime pattern is currently the most scalable approach: it does not intrude on existing agent logic and does not consume the context window, yet it provides independent, continuous compliance monitoring. Our
score_enforcer.pyis evolving in this direction.
This article is based on a comprehensive Loop Engineering audit conducted in the Junze Zhiku production environment on June 14, 2026. All code references come from the corresponding projects' public source code or official documentation. Special thanks to the DeepSeek team for providing an excellent model; it was precisely its "failures" that allowed us to dig deeper into the underlying problems of this field.
More in Evidence
- A Reality Check on Decision Models: Why They Seem Miraculous Online but We Measured Only 54%: A Full Comparison of JEV / LAYA / KEV / CLM-8B and a Deployment Formula
- The "Non-Text-Generating Model": Jev and the New System One Category, and How Agent Architecture Changes When AI Only Answers Multiple Choice
- WeChat Open Source WeMM-Embedding Deep Dive: The Multimodal Embedding Model Topping MMEB-v2, Can It Run on Your Mac?
- A Source-Level Architectural Dissection of DeepSeek Harness: How an Everything-Is-a-Plugin Agent Framework Is Built