Loop Engineering: A Deep Retrospective on a Field Failure, When the Design Documents Cannot Be Executed
import { ArticleLayout } from '@/layouts/article-layout' export default ArticleLayout
Loop Engineering: A Deep Retrospective on a Field Failure
When the Design Documents Cannot Be Executed
On June 7, 2026, OpenClaw founder Peter Steinberger posted: "You shouldn't be prompting coding agents anymore. You should be designing loops that prompt your agents." The same day, Anthropic Claude Code lead Boris Cherny said publicly: "I don't prompt Claude anymore. I have loops running. My job is to write loops."
On June 8, Google engineering director Addy Osmani published a long article formally naming this methodology Loop Engineering.
On June 13, we spent an entire day trying to build a complete system from scratch.
Result: failure.
What We Actually Built
36 files. 7 modules. 10 cron jobs. 8 Loops. 1 Dashboard.
On paper, this is a complete Loop Engineering system:
- A three-layer architecture (Main → Looper → coder-*)
- A seven-phase inner loop (Observe → Plan → Execute → Verify → Diagnose → Adjust → Retry)
- A cap of 30 self-corrections + a four-level escalating strategy
- A five-dimensional meaning evaluation model (only ≥70 points gets converted into a Loop)
The Truth: Nothing Happened
When we used openclaw cron run to trigger a deliberately failing test Loop, the file was not created, and STATE.json was completely empty.
Root cause: there was no Loop Runtime Engine.
The YAML frontmatter of LOOP.md defined executor, checker, inner_loop, and escalation. But after cron triggered, there was no code at all to parse and execute those definitions.
Here is the gap between industry standards and our reality:
| Industry Standard | Us |
|---|---|
| Maker ≠ Checker | ❌ At the design stage, both were the looper |
| Separate verifier model | ❌ Written, but no code executed it |
| /goal: run until condition met | ❌ The inner loop never happened |
| State on disk, not in context | ✅ Designed correctly |
Three Fatal Mistakes
1. A file ≠ a system. 36 design documents are not the same as a runnable system.
2. sessions_spawn ≠ a Loop. We used sessions_spawn to bypass the inner loop and complete all our "deliverables". This is not Loop Engineering; it is scheduling multiple Agents.
3. Blind spots are only exposed at runtime. The Maker/Checker Split error was completely undetected at the design stage, because the system had never been executed.
Next Steps
Core insight: the hard part of Loop Engineering was never writing LOOP.md, but writing the engine that can execute LOOP.md.
Three paths: a Python Loop Runner (a 200-line build of our own), Claude Code /goal (a native mechanism), or a hybrid approach.
More in Evidence
- A Reality Check on Decision Models: Why They Seem Miraculous Online but We Measured Only 54%: A Full Comparison of JEV / LAYA / KEV / CLM-8B and a Deployment Formula
- The "Non-Text-Generating Model": Jev and the New System One Category, and How Agent Architecture Changes When AI Only Answers Multiple Choice
- WeChat Open Source WeMM-Embedding Deep Dive: The Multimodal Embedding Model Topping MMEB-v2, Can It Run on Your Mac?
- A Source-Level Architectural Dissection of DeepSeek Harness: How an Everything-Is-a-Plugin Agent Framework Is Built