Agentic Research

Technical Argument for the OpenClaw Loop Engineering Refactor: A Complete Plan from Prompt-Driven to Architecturally Enforced

2026/06/1444 min readBryan Chan閱讀中文原文
TopicsOpenClawArchitectureLoop EngineeringAgent Architecture

Background: We have run 10 Auto-Healing Loops in the production environment, gone through a complete audit of 204 sessions, and compared 14 AI Agent architecture approaches on the market. Now it is time to turn those findings into a concrete system refactor plan. Structure of this article: current-state diagnosis → 8 refactor hypotheses → a comparison of technology options across 8 architecture nodes → a complete refactor roadmap


1. Current-State Diagnosis: Six Structural Defects in the Current Architecture

1.1 Execution and Verification Are Not Separated

ProblemEvidenceImpact
The main agent both executes and verifies itselfsystem-health skipped the OMLX check 5 times in a row, with a verification spawn rate of 0/5A critical service outage went undetected for 20 hours
The looper both edits and accepts its own codefull_repair session: 55 tool calls, 0 checker spawnsChanges went unreviewed by an independent party
The verify instruction at the prompt level is ignored6 sessions, 0 checker spawnsPlain-text rules fail 100% of the time in an LLM

1.2 The Dual Cron System Is Chaotic

ProblemEvidence
The main OpenClaw cron has 11 jobs (including loop jobs)system-health + daily-triage are bound to the main agent
The looper has its own separate cron system (different cron ID prefixes)6 jobs such as pr-review and email-watch go through the looper
The two systems have no unified managementThere is no single view of all loop schedules

1.3 Workspace Paths Are Chaotic

AgentWorkspaceActual path when reading LOOP.md
main~/.openclaw/workspace/~/.openclaw/workspace/loops/... (old version!)
looper~/Desktop/UltraClaw_Project/loop-engineering/Correct

1.4 No Methodology Enforcement Mechanism

  • The seven-phase inner loop of LOOP.md exists only in prompt text
  • There is no architecture-level completion gate (an equivalent of the Stop Hook)
  • compliance is not tracked, audited, or enforced

1.5 The Delivery Mechanism Is Misconfigured

  • 6 looper cron jobs try to send Feishu messages directly (LarkClient appId/appSecret missing)
  • The looper should send via sessions_send to the main agent, rather than sending directly

1.6 State Management Is Fragmented

  • Each loop has its own STATE.json
  • There is no global state aggregation (state/aggregator.py exists but is not used by cron)
  • It is impossible to query "the current state of all loops" from a single node

2. Refactor Hypotheses

Based on market research and the production audit, we propose the following 8 refactor hypotheses:

#HypothesisSource of InspirationVerification Method
H1Promoting verify from prompt to architectural enforcement can raise compliance from ~0% to ~100%Claude Code Stop HookCompare compliance_score before and after the refactor
H2Role isolation (execute agent ≠ verify agent) is more than 10x more reliable than a prompt suggestionMetaGPT QA EngineerA/B testing
H3A unified state management layer can improve audit efficiency by 5xVIGIL EmoBankCompare audit time
H4An external reflection layer (an independent cron scan) can catch failures that prompt-level measures cannot preventVIGIL Out-of-bandCompare failure detection rates
H5Using the Race Pattern for critical decisions (multiple models judging in parallel) can cut the misjudgment rate by 50%+Cursor Race PatternAccuracy of critical decisions
H6Turning the inner loop into an engine (inner_loop_runner.py) can raise the retry execution rate from 0% to 100%OpenHands Loop RecoveryNumber of retry executions
H7A unified Delivery Bus can eliminate 100% of cross-agent communication errorsNo specific source, problem-drivenCommunication error rate
H8Separating Planning from Execution in cron can raise scheduling flexibility by 3xAider Architect-EditorCost of schedule changes

3. Technology Options for Architecture Nodes

Node 1: Agent Role Isolation Mechanism

Problem: How do you ensure the execute agent never verifies its own work?

OptionDescriptionEnforcementImplementation ComplexityChange to Existing System
A: Config-level (current)Set roles in the agent config and suggest behavior in the prompt🟢 NoneLowNone
B: Tool-gatingIntercept at the Gateway layer: a verify agent cannot exec, an execute agent cannot spawn a checker🟠 ModerateMediumGateway plugin
C: Stop HookBorrowing from Claude Code: forcibly run an external verification script before a loop completes, and block if it does not pass🔴 StrongMediumModify the Gateway turn lifecycle
D: Two-phase cronSplit each loop into two independent cron jobs (execute + verify) with different agents🔴 StrongLowRewrite the cron job definitions

Recommendation: C + D combined. Stop Hook as the real-time line of defense, two-phase cron as the structural guarantee.


Node 2: Verification Mechanism (Checker)

Problem: Who verifies? How? And what happens when verification fails?

OptionDescriptionVerification SignalDegree of Autonomy
A: Self-verify (current)The agent checks its own result in the promptSoft (textual judgment)❌ Unstable
B: Dedicated checker agentSpawn coder-deepseek for an independent review after each loop completesSoft (agent judgment)🟠 Requires the agent to spawn proactively
C: Stop Hook scriptA Python/Shell script checks the STATE.json schema + key metricsHard (structured check)🔴 Externally enforced
D: Race PatternFor critical decisions (such as OMLX severity), have 2+ agents judge at once and take the majoritySoft (majority vote)🔴 Statistical guarantee
E: VIGIL-style supervisorAn independent process periodically scans all loop STATE.json files and checks complianceHard (structured check)🔴 External and continuous

Recommendation: C + E as the mainstay, with D as an enhancement for critical decisions.


Node 3: Cron / Scheduling System

Problem: How do you uniformly manage the schedules of 10+ loops? Should Planning and Execution be separated?

OptionDescriptionProsCons
A: OpenClaw built-in (current)Use the OpenClaw cron APIDeep integration with the GatewayDual systems, agent binding limitations
B: Native crontabTrigger with the system crontabStable, independent of the GatewayNo built-in retry, no state tracking
C: Scheduler AgentAn independent scheduler agent manages all schedulesSingle source of truth, flexibleSingle point of failure
D: Plan-Execute separationManage the Planning cron (when to run) and the Execution cron (who runs) separatelyClear responsibilities, easy to changeTwo-layer scheduling coordination

Recommendation: C + D. Establish a loop-scheduler agent to manage schedules uniformly, with Plan (the cron definition) and Execute (the actual trigger) separated.


Node 4: State Management

Problem: How do you uniformly track the execution state, compliance, and history of 10 loops?

OptionDescriptionQuery CapabilityChange TrackingComplexity
A: STATE.json per loop (current)One JSON file per loopManualgit diffLow
B: Centralized SQLiteAll loop state written to a single DBSQL queryBuilt-inMedium
C: Event-sourcingAll state changes appended as events; state is rebuilt from eventsTime travelCompleteHigh
D: Global aggregatorstate/aggregator.py periodically aggregates all STATE.json files → writes global-state.jsonSingle fileSnapshot comparisonLow

Recommendation: B + D. SQLite as the mainstay, with global-state.json as a quick-view interface.


Node 5: Inner Loop Engine (Retry)

Problem: How do you ensure a loop actually retries on failure, rather than reporting once and stopping?

OptionDescriptionTrigger MethodEnforcement
A: Prompt-based (current)Describe the retry strategy in LOOP.md and let the agent decideAgent self-discipline🟢 None
B: Cron-level retrySet retry on failure on the cron jobGateway🟠 Configurable
C: Engine-levelinner_loop_runner.py forcibly runs the retry loopScript🔴 Strong
D: State-machineDefine legal state transitions; illegal transitions raise errors (VIGIL-style)State machine🔴 Strongest

Recommendation: C. Make inner_loop_runner.py the standard engine, which every loop must run through.


Node 6: Delivery / Communication Bus

Problem: How do agents communicate? How do you ensure an escalation reaches the right destination?

OptionDescriptionReliabilityComplexity
A: Direct channel (current)The agent sends a message directly to a specific channel🟡 Depends on the agent configLow
B: sessions_send busAll agents send to main via sessions_send, and main forwards everything🟠 Depends on mainLow
C: Message queueRedis/RabbitMQ as middleware🔴 HighHigh
D: Escalation engineAn independent escalation agent that routes automatically by severity + target🔴 HighMedium

Recommendation: B + D. sessions_send for everyday communication, an escalation engine to guarantee delivery of critical alerts.


Node 7: Context Management

Problem: How do long-running loops avoid context bloat? How do they retain key decisions?

OptionDescriptionApplicable Scenario
A: Session isolation (current)Create a new session for each cron runShort-term tasks
B: STATE.json persistence (current)Key state is written to STATE.json and read on the next runCross-session memory
C: Claude Code-style compressionA five-layer compression pipeline that preserves key decisionsLong-running sessions
D: Aider-style RepoMapInject only relevant context, not all historyLarge-scale tasks

Recommendation: Keep A + B (already sufficient); for loops that need long-term context, add C's key-decision retention mechanism.


Node 8: Multi-Model Strategy

Problem: Do different loops need different models? Do critical decisions need multi-model verification?

OptionDescriptionCostReliability
A: Fixed assignment (current)Each agent has a fixed modelLowDepends on the model
B: Dynamic routerSelect the model dynamically by task complexityMedium🟠
C: Race PatternFor critical decisions, call multiple models in parallel and take the bestHigh🔴
D: Architect-EditorA strong reasoning model plans, a fast model executesMedium🟠

Recommendation: A as the baseline, C as an enhancement for OMLX-level critical decisions, and D as an option for complex loops (such as rnd-discovery).


4. Refactor Roadmap

Phase 0: Non-Destructive Prerequisites (this week, 0 risk)

#ActionDependencyEstimate
P0.1Unified state management layer: SQLite + global-state.jsonNone2h
P0.2Deploy the loop-scheduler agent: a unified cron viewNone3h
P0.3Fix all delivery settings + escalation engineNone2h
P0.4Clean up the old workspace directory (partially done)None0.5h

Phase 1: Architecture Hardening (next week, low risk)

#ActionDependencyEstimate
P1.1Implement the Stop Hook: modify the Gateway turn lifecycleP04h
P1.2Split the loop cron into two layers, Plan + ExecuteP0.1, P0.23h
P1.3Deploy two-phase cron (execute + verify separated)P1.24h
P1.4Implement the inner_loop_runner.py engineP0.15h

Phase 2: Intelligence Enhancement (this month, moderate risk)

#ActionDependencyEstimate
P2.1A VIGIL-style external reflection layer (an extension of score_enforcer)P0.16h
P2.2Race Pattern for OMLX-level critical decisionsP1.44h
P2.3Dynamic model router for complex loopsP0.25h
P2.4Claude Code-style context compactionP1.48h

Phase 3: Ecosystem Refinement (next month, low risk)

#ActionDependencyEstimate
P3.1A loop template engine: create a new loop with one commandP1.3, P1.46h
P3.2Dashboard v2: a global loop health + compliance viewP0.18h
P3.3An automated regression test suiteP1.46h
P3.4Documentation + community releaseP3.1-3.34h

5. Expected Outcomes

Quantitative Goals

MetricCurrentPhase 0Phase 1Phase 2
Verify execution rate0%0%100%100%
Compliance score~2/74/76/77/7
Unified cron managementDual systemSingleSingleSingle
Delivery error rate6/10 jobs0/100/100/10
Critical decision misjudgment rate (OMLX)100% (5/5)100%100%<20%
Audit efficiencyManual, 204 sessionsAutomatedAutomatedAutomated + predictive

Qualitative Goals

  • From "prompt prayer" to "architectural guarantee": no longer relying on the agent to follow rules out of self-discipline
  • From "black-box execution" to "transparent audit": every step of every loop is traceable
  • From "passive repair" to "proactive prevention": the external reflection layer detects patterns before problems occur
  • From "single-point decision" to "multi-party verification": critical judgments execute only after multi-party confirmation

6. Risks and Mitigation

RiskImpactMitigation
The Stop Hook over-blocks the normal flowService interruptionAn 8-block cap + gradual rollout
SQLite concurrent write conflictsState lossWAL mode + a write queue
The Race Pattern increases API costBudget overrunUse only for critical decisions (<5% of decision points)
Cron job interruption during the architecture changeMonitoring gapsPhase 0 prerequisites do not affect existing cron
An infinite loop in inner_loop_runnerResource exhaustionA hard limit of max_retries=30

This article is a technical argument draft for the OpenClaw Loop Engineering refactor. All hypotheses need to be discussed and prioritized before being turned into a concrete implementation plan.