Technical Argument for the OpenClaw Loop Engineering Refactor: A Complete Plan from Prompt-Driven to Architecturally Enforced
Background: We have run 10 Auto-Healing Loops in the production environment, gone through a complete audit of 204 sessions, and compared 14 AI Agent architecture approaches on the market. Now it is time to turn those findings into a concrete system refactor plan. Structure of this article: current-state diagnosis → 8 refactor hypotheses → a comparison of technology options across 8 architecture nodes → a complete refactor roadmap
1. Current-State Diagnosis: Six Structural Defects in the Current Architecture
1.1 Execution and Verification Are Not Separated
| Problem | Evidence | Impact |
|---|---|---|
| The main agent both executes and verifies itself | system-health skipped the OMLX check 5 times in a row, with a verification spawn rate of 0/5 | A critical service outage went undetected for 20 hours |
| The looper both edits and accepts its own code | full_repair session: 55 tool calls, 0 checker spawns | Changes went unreviewed by an independent party |
| The verify instruction at the prompt level is ignored | 6 sessions, 0 checker spawns | Plain-text rules fail 100% of the time in an LLM |
1.2 The Dual Cron System Is Chaotic
| Problem | Evidence |
|---|---|
| The main OpenClaw cron has 11 jobs (including loop jobs) | system-health + daily-triage are bound to the main agent |
| The looper has its own separate cron system (different cron ID prefixes) | 6 jobs such as pr-review and email-watch go through the looper |
| The two systems have no unified management | There is no single view of all loop schedules |
1.3 Workspace Paths Are Chaotic
| Agent | Workspace | Actual path when reading LOOP.md |
|---|---|---|
| main | ~/.openclaw/workspace/ | ~/.openclaw/workspace/loops/... (old version!) |
| looper | ~/Desktop/UltraClaw_Project/loop-engineering/ | Correct |
1.4 No Methodology Enforcement Mechanism
- The seven-phase inner loop of LOOP.md exists only in prompt text
- There is no architecture-level completion gate (an equivalent of the Stop Hook)
- compliance is not tracked, audited, or enforced
1.5 The Delivery Mechanism Is Misconfigured
- 6 looper cron jobs try to send Feishu messages directly (LarkClient appId/appSecret missing)
- The looper should send via sessions_send to the main agent, rather than sending directly
1.6 State Management Is Fragmented
- Each loop has its own STATE.json
- There is no global state aggregation (
state/aggregator.pyexists but is not used by cron) - It is impossible to query "the current state of all loops" from a single node
2. Refactor Hypotheses
Based on market research and the production audit, we propose the following 8 refactor hypotheses:
| # | Hypothesis | Source of Inspiration | Verification Method |
|---|---|---|---|
| H1 | Promoting verify from prompt to architectural enforcement can raise compliance from ~0% to ~100% | Claude Code Stop Hook | Compare compliance_score before and after the refactor |
| H2 | Role isolation (execute agent ≠ verify agent) is more than 10x more reliable than a prompt suggestion | MetaGPT QA Engineer | A/B testing |
| H3 | A unified state management layer can improve audit efficiency by 5x | VIGIL EmoBank | Compare audit time |
| H4 | An external reflection layer (an independent cron scan) can catch failures that prompt-level measures cannot prevent | VIGIL Out-of-band | Compare failure detection rates |
| H5 | Using the Race Pattern for critical decisions (multiple models judging in parallel) can cut the misjudgment rate by 50%+ | Cursor Race Pattern | Accuracy of critical decisions |
| H6 | Turning the inner loop into an engine (inner_loop_runner.py) can raise the retry execution rate from 0% to 100% | OpenHands Loop Recovery | Number of retry executions |
| H7 | A unified Delivery Bus can eliminate 100% of cross-agent communication errors | No specific source, problem-driven | Communication error rate |
| H8 | Separating Planning from Execution in cron can raise scheduling flexibility by 3x | Aider Architect-Editor | Cost of schedule changes |
3. Technology Options for Architecture Nodes
Node 1: Agent Role Isolation Mechanism
Problem: How do you ensure the execute agent never verifies its own work?
| Option | Description | Enforcement | Implementation Complexity | Change to Existing System |
|---|---|---|---|---|
| A: Config-level (current) | Set roles in the agent config and suggest behavior in the prompt | 🟢 None | Low | None |
| B: Tool-gating | Intercept at the Gateway layer: a verify agent cannot exec, an execute agent cannot spawn a checker | 🟠 Moderate | Medium | Gateway plugin |
| C: Stop Hook | Borrowing from Claude Code: forcibly run an external verification script before a loop completes, and block if it does not pass | 🔴 Strong | Medium | Modify the Gateway turn lifecycle |
| D: Two-phase cron | Split each loop into two independent cron jobs (execute + verify) with different agents | 🔴 Strong | Low | Rewrite the cron job definitions |
Recommendation: C + D combined. Stop Hook as the real-time line of defense, two-phase cron as the structural guarantee.
Node 2: Verification Mechanism (Checker)
Problem: Who verifies? How? And what happens when verification fails?
| Option | Description | Verification Signal | Degree of Autonomy |
|---|---|---|---|
| A: Self-verify (current) | The agent checks its own result in the prompt | Soft (textual judgment) | ❌ Unstable |
| B: Dedicated checker agent | Spawn coder-deepseek for an independent review after each loop completes | Soft (agent judgment) | 🟠 Requires the agent to spawn proactively |
| C: Stop Hook script | A Python/Shell script checks the STATE.json schema + key metrics | Hard (structured check) | 🔴 Externally enforced |
| D: Race Pattern | For critical decisions (such as OMLX severity), have 2+ agents judge at once and take the majority | Soft (majority vote) | 🔴 Statistical guarantee |
| E: VIGIL-style supervisor | An independent process periodically scans all loop STATE.json files and checks compliance | Hard (structured check) | 🔴 External and continuous |
Recommendation: C + E as the mainstay, with D as an enhancement for critical decisions.
Node 3: Cron / Scheduling System
Problem: How do you uniformly manage the schedules of 10+ loops? Should Planning and Execution be separated?
| Option | Description | Pros | Cons |
|---|---|---|---|
| A: OpenClaw built-in (current) | Use the OpenClaw cron API | Deep integration with the Gateway | Dual systems, agent binding limitations |
| B: Native crontab | Trigger with the system crontab | Stable, independent of the Gateway | No built-in retry, no state tracking |
| C: Scheduler Agent | An independent scheduler agent manages all schedules | Single source of truth, flexible | Single point of failure |
| D: Plan-Execute separation | Manage the Planning cron (when to run) and the Execution cron (who runs) separately | Clear responsibilities, easy to change | Two-layer scheduling coordination |
Recommendation: C + D. Establish a loop-scheduler agent to manage schedules uniformly, with Plan (the cron definition) and Execute (the actual trigger) separated.
Node 4: State Management
Problem: How do you uniformly track the execution state, compliance, and history of 10 loops?
| Option | Description | Query Capability | Change Tracking | Complexity |
|---|---|---|---|---|
| A: STATE.json per loop (current) | One JSON file per loop | Manual | git diff | Low |
| B: Centralized SQLite | All loop state written to a single DB | SQL query | Built-in | Medium |
| C: Event-sourcing | All state changes appended as events; state is rebuilt from events | Time travel | Complete | High |
| D: Global aggregator | state/aggregator.py periodically aggregates all STATE.json files → writes global-state.json | Single file | Snapshot comparison | Low |
Recommendation: B + D. SQLite as the mainstay, with global-state.json as a quick-view interface.
Node 5: Inner Loop Engine (Retry)
Problem: How do you ensure a loop actually retries on failure, rather than reporting once and stopping?
| Option | Description | Trigger Method | Enforcement |
|---|---|---|---|
| A: Prompt-based (current) | Describe the retry strategy in LOOP.md and let the agent decide | Agent self-discipline | 🟢 None |
| B: Cron-level retry | Set retry on failure on the cron job | Gateway | 🟠 Configurable |
| C: Engine-level | inner_loop_runner.py forcibly runs the retry loop | Script | 🔴 Strong |
| D: State-machine | Define legal state transitions; illegal transitions raise errors (VIGIL-style) | State machine | 🔴 Strongest |
Recommendation: C. Make inner_loop_runner.py the standard engine, which every loop must run through.
Node 6: Delivery / Communication Bus
Problem: How do agents communicate? How do you ensure an escalation reaches the right destination?
| Option | Description | Reliability | Complexity |
|---|---|---|---|
| A: Direct channel (current) | The agent sends a message directly to a specific channel | 🟡 Depends on the agent config | Low |
| B: sessions_send bus | All agents send to main via sessions_send, and main forwards everything | 🟠 Depends on main | Low |
| C: Message queue | Redis/RabbitMQ as middleware | 🔴 High | High |
| D: Escalation engine | An independent escalation agent that routes automatically by severity + target | 🔴 High | Medium |
Recommendation: B + D. sessions_send for everyday communication, an escalation engine to guarantee delivery of critical alerts.
Node 7: Context Management
Problem: How do long-running loops avoid context bloat? How do they retain key decisions?
| Option | Description | Applicable Scenario |
|---|---|---|
| A: Session isolation (current) | Create a new session for each cron run | Short-term tasks |
| B: STATE.json persistence (current) | Key state is written to STATE.json and read on the next run | Cross-session memory |
| C: Claude Code-style compression | A five-layer compression pipeline that preserves key decisions | Long-running sessions |
| D: Aider-style RepoMap | Inject only relevant context, not all history | Large-scale tasks |
Recommendation: Keep A + B (already sufficient); for loops that need long-term context, add C's key-decision retention mechanism.
Node 8: Multi-Model Strategy
Problem: Do different loops need different models? Do critical decisions need multi-model verification?
| Option | Description | Cost | Reliability |
|---|---|---|---|
| A: Fixed assignment (current) | Each agent has a fixed model | Low | Depends on the model |
| B: Dynamic router | Select the model dynamically by task complexity | Medium | 🟠 |
| C: Race Pattern | For critical decisions, call multiple models in parallel and take the best | High | 🔴 |
| D: Architect-Editor | A strong reasoning model plans, a fast model executes | Medium | 🟠 |
Recommendation: A as the baseline, C as an enhancement for OMLX-level critical decisions, and D as an option for complex loops (such as rnd-discovery).
4. Refactor Roadmap
Phase 0: Non-Destructive Prerequisites (this week, 0 risk)
| # | Action | Dependency | Estimate |
|---|---|---|---|
| P0.1 | Unified state management layer: SQLite + global-state.json | None | 2h |
| P0.2 | Deploy the loop-scheduler agent: a unified cron view | None | 3h |
| P0.3 | Fix all delivery settings + escalation engine | None | 2h |
| P0.4 | Clean up the old workspace directory (partially done) | None | 0.5h |
Phase 1: Architecture Hardening (next week, low risk)
| # | Action | Dependency | Estimate |
|---|---|---|---|
| P1.1 | Implement the Stop Hook: modify the Gateway turn lifecycle | P0 | 4h |
| P1.2 | Split the loop cron into two layers, Plan + Execute | P0.1, P0.2 | 3h |
| P1.3 | Deploy two-phase cron (execute + verify separated) | P1.2 | 4h |
| P1.4 | Implement the inner_loop_runner.py engine | P0.1 | 5h |
Phase 2: Intelligence Enhancement (this month, moderate risk)
| # | Action | Dependency | Estimate |
|---|---|---|---|
| P2.1 | A VIGIL-style external reflection layer (an extension of score_enforcer) | P0.1 | 6h |
| P2.2 | Race Pattern for OMLX-level critical decisions | P1.4 | 4h |
| P2.3 | Dynamic model router for complex loops | P0.2 | 5h |
| P2.4 | Claude Code-style context compaction | P1.4 | 8h |
Phase 3: Ecosystem Refinement (next month, low risk)
| # | Action | Dependency | Estimate |
|---|---|---|---|
| P3.1 | A loop template engine: create a new loop with one command | P1.3, P1.4 | 6h |
| P3.2 | Dashboard v2: a global loop health + compliance view | P0.1 | 8h |
| P3.3 | An automated regression test suite | P1.4 | 6h |
| P3.4 | Documentation + community release | P3.1-3.3 | 4h |
5. Expected Outcomes
Quantitative Goals
| Metric | Current | Phase 0 | Phase 1 | Phase 2 |
|---|---|---|---|---|
| Verify execution rate | 0% | 0% | 100% | 100% |
| Compliance score | ~2/7 | 4/7 | 6/7 | 7/7 |
| Unified cron management | Dual system | Single | Single | Single |
| Delivery error rate | 6/10 jobs | 0/10 | 0/10 | 0/10 |
| Critical decision misjudgment rate (OMLX) | 100% (5/5) | 100% | 100% | <20% |
| Audit efficiency | Manual, 204 sessions | Automated | Automated | Automated + predictive |
Qualitative Goals
- From "prompt prayer" to "architectural guarantee": no longer relying on the agent to follow rules out of self-discipline
- From "black-box execution" to "transparent audit": every step of every loop is traceable
- From "passive repair" to "proactive prevention": the external reflection layer detects patterns before problems occur
- From "single-point decision" to "multi-party verification": critical judgments execute only after multi-party confirmation
6. Risks and Mitigation
| Risk | Impact | Mitigation |
|---|---|---|
| The Stop Hook over-blocks the normal flow | Service interruption | An 8-block cap + gradual rollout |
| SQLite concurrent write conflicts | State loss | WAL mode + a write queue |
| The Race Pattern increases API cost | Budget overrun | Use only for critical decisions (<5% of decision points) |
| Cron job interruption during the architecture change | Monitoring gaps | Phase 0 prerequisites do not affect existing cron |
| An infinite loop in inner_loop_runner | Resource exhaustion | A hard limit of max_retries=30 |
This article is a technical argument draft for the OpenClaw Loop Engineering refactor. All hypotheses need to be discussed and prioritized before being turned into a concrete implementation plan.
More in Evidence
- A Reality Check on Decision Models: Why They Seem Miraculous Online but We Measured Only 54%: A Full Comparison of JEV / LAYA / KEV / CLM-8B and a Deployment Formula
- The "Non-Text-Generating Model": Jev and the New System One Category, and How Agent Architecture Changes When AI Only Answers Multiple Choice
- WeChat Open Source WeMM-Embedding Deep Dive: The Multimodal Embedding Model Topping MMEB-v2, Can It Run on Your Mac?
- A Source-Level Architectural Dissection of DeepSeek Harness: How an Everything-Is-a-Plugin Agent Framework Is Built