Deep Research on Prime Agent: PrimeIntellect's Self-Improving RLM Agent and Our Field Test
The One-Sentence Version
Prime Intellect (a distributed AI training company) open-sourced Prime Agent on August 6, 2026: a coding/research agent designed for long-task autonomous operation. MIT licensed, one-command install, with a TUI based on a fork of pi-mono. Within a week of launch, it was covered by multiple media outlets.
Official results: Opus 5 + Prime Agent scored 95.5% on ARC-AGI-3, surpassing the human expert baseline of 95.4%. In our field test, it wrote working Rust emulators for the SEGA Genesis and Game Boy Color without reference code.
Two Core Abstractions
Prime Agent's design revolves around two paper-grade abstractions, which is also its fundamental divergence from traditional agent harnesses (fixed tool schemas + context compression).
1. Recursive Language Model (RLM)
Context is a variable, and a subagent is a function call.
A traditional harness gives the model a pile of fixed-schema tools (read_file, run_command...) and then uses context compression to "help" the model save tokens. RLM does the opposite:
- The model has only one tool: a persistent IPython kernel
- Context is handled as a variable (prompt-as-a-variable); message history is stored in variables, and the model can write code against its own context
- A subagent is just a function call:
rlm("subtask")spawns a complete subagent (with its own model, kernel, and history) and returns non-blockingly, with the result delivered asynchronously
The effect: in theory it supports infinitely long sessions without losing history, because history is not "compressed away" but "referenced programmatically".
2. Continual Harness (the Self-Evolving Shell)
The harness state H = (prompts, sub-agents, skills, memory), and the Agent can CRUD it.
The /refine command reads the agent's own execution traces, extracts lessons, and makes minimal, evidence-backed updates to the harness state:
- Update types: supplementing prompts, memory, skill descriptions, subagent specs
- The base system prompt is never mutable (the safety anchor)
- Every update records a snapshot and can be rolled back by ID
- Planning runs in the background without blocking the conversation
This is the engineering realization of "an agent improving itself": not changing weights, but changing its own operating environment.
Other Engineering Features
| Feature | Description |
|---|---|
| Daemon-backed sessions | The agent keeps running after the terminal disconnects, and you can reattach at any time |
| Direct inter-agent communication | Running agents can discover each other, exchange messages, and work together |
| Heartbeats & Schedules | Periodic/scheduled re-entry into a session |
| Persistent Goals | /goal keeps a goal alive across turns until it is complete |
| Bounded Autonomous Mode | /autonomous executes autonomously within a turn/token/time budget + a custom quality gate |
| Skills = executable Python packages | A built-in skill creator distills workflows into skills |
| Headless mode | JSON / RPC modes for automation integration |
🪞 Comparison with UltraClaw Infrastructure (the Focus of This Article)
On our own OpenClaw system, we assembled nearly identical capabilities using external infrastructure. Here is the comparison table:
| Prime Agent (harness-native) | UltraClaw (OpenClaw + plugin layer) |
|---|---|
Continual Harness /refine | The three-piece suite of agent-evolver + skill-curator + SkillOpt |
| RLM subagent function calls | sessions_spawn + subagents orchestration |
| Daemon + reattach | The OpenClaw Gateway resident process + sessions |
| Heartbeats / Schedules | OpenClaw cron + heartbeat (the same mechanisms) |
| Skills = Python packages | The SKILL.md skill system + SkillOpt nightly optimization |
| Snapshot + rollback | Our verify.py + Gate JSON validation |
| Inter-agent communication | sessions_send + multi-agent routing |
Conclusion: the direction is exactly the same, which proves we are on the frontier path. The difference is at the implementation level:
- Prime Agent builds self-evolution into the harness kernel (
/refineis a first-class citizen); ours is a plugin layer stacked on top (three independent skills + external cron) - Prime Agent's RLM solves long-task memory with "programmatic context"; we use a three-layer file system of vector memory + daily log + session transcript
- Our advantages: multi-channel integration (Feishu/WeChat/WhatsApp), a business-scenario skill library (400+ skills such as AK-SDD Hong Kong stock research), and a compliance audit layer. These are vertical accumulations a general-purpose harness will not build
One interesting observation: Prime Agent's /refine never touches the base system prompt, and the pit we fell into (Lesson #76-77: model switching causing behavioral regression) precisely demonstrates the importance of an immutable anchor. Different teams independently converging on the same design suggests this is the correct answer.
🔬 Field Test Record (Mac Studio, early hours of 2026-08-13)
Installation Audit
The official install script is 1,620 lines; we audited it before installing (rather than blindly running curl | sh):
- SHA-256 verification of the release package ✅
- Sudo is needed only when Node.js is missing ✅
- No suspicious outbound connections or persistence behavior ✅
Installation result: prime-agent v0.7.2, 192 dependency packages, checksum verification passed.
Connecting to Our Own API Relay Layer (Sub2API)
Prime Agent supports adding any OpenAI-compatible provider via ~/.prime/agent/models.json:
{
"providers": {
"sub2api": {
"baseUrl": "http://localhost:8080/v1",
"api": "openai-completions",
"apiKey": "sk-...",
"compat": {
"supportsDeveloperRole": false,
"supportsReasoningEffort": false
},
"models": [
{ "id": "deepseek-v4-pro", "reasoning": false, "contextWindow": 128000 }
]
}
}
}
Pitfall: Cloudflare Error 1010
The first test returned 403 Your request was blocked. The troubleshooting process:
- curl direct to the API → 200 ✅ (ruling out a key issue)
- curl simulating stream/tools/UA → all 200 ✅
- A local capture server grabbed prime-agent's complete request → replayed it to the production domain → 403, error code: 1010
Root cause: Cloudflare WAF's rule 1010 (blocking by browser signature). Prime Agent uses the OpenAI JS SDK, with a User-Agent of OpenAI/JS 6.47.0 plus a string of X-Stainless-* fingerprint headers, which the WAF identified as bot SDK traffic and blocked.
Solution: switch to a local endpoint (http://localhost:8080/v1) to bypass Cloudflare. Sub2API runs on the same machine anyway, so latency is lower.
Field Test Results
Test 1 (basic connectivity):
> Reply with exactly: PRIME_AGENT_TEST_OK
< PRIME_AGENT_TEST_OK ✅
Test 2 (a real coding task): asked to create a memoized version of fibonacci and verify fib(50):
- The Agent wrote files, executed, and verified through the IPython kernel
- It produced
fibonacci.py(an@lru_cacheimplementation, with boundary checks and a docstring) - It correctly returned
fib(50) = 12586269025✅
The whole process took about 40 seconds, passed on the first try, and needed no human intervention.
⚠️ Risks and Notes
- It is not a security sandbox: the official README explicitly warns that it executes model-generated Python with the user's privileges. The worker/kernel processes have lifecycle isolation, but that does not constitute a security boundary. Use an external sandbox to run untrusted code.
- Cost: it is designed for frontier models, and the token consumption of long tasks + recursive subagents is considerable.
- A young ecosystem: it was open-sourced only on August 6; the documentation is complete, but community examples are still few.
What It Means for Us
- Borrow
/refine's snapshot + rollback mechanism: our SkillOpt already has a gate, but lacks the fine-grained control of "rolling back a single update by ID" - RLM's programmatic context is worth tracking: when OpenClaw supports it, long-task memory may no longer need a three-layer file system
- Cloudflare WAF will falsely flag legitimate agent traffic: this is a new pitfall at the infrastructure layer, worth adding to the operations manual
- Self-improvement is not the future, it is now: Prime Agent, our seven-piece suite, and SkillOpt are all converging on this direction in 2026
Prime Agent is worth tracking continuously. Our next step: run a real long task in an isolated environment (for example, a complete AK-SDD research workflow) and test whether the Continual Harness's /refine can produce valuable self-improvement.
This article is based on a field test of Prime Agent v0.7.2, 2026-08-13.
More in Evidence
- A Reality Check on Decision Models: Why They Seem Miraculous Online but We Measured Only 54%: A Full Comparison of JEV / LAYA / KEV / CLM-8B and a Deployment Formula
- The "Non-Text-Generating Model": Jev and the New System One Category, and How Agent Architecture Changes When AI Only Answers Multiple Choice
- WeChat Open Source WeMM-Embedding Deep Dive: The Multimodal Embedding Model Topping MMEB-v2, Can It Run on Your Mac?
- A Source-Level Architectural Dissection of DeepSeek Harness: How an Everything-Is-a-Plugin Agent Framework Is Built