Agentic Research

Deep Research on Prime Agent: PrimeIntellect's Self-Improving RLM Agent and Our Field Test

2026/08/1329 min readBryan Chan閱讀中文原文
TopicsPrime AgentPrimeIntellectRLMAgent FrameworkOpen Source

The One-Sentence Version

Prime Intellect (a distributed AI training company) open-sourced Prime Agent on August 6, 2026: a coding/research agent designed for long-task autonomous operation. MIT licensed, one-command install, with a TUI based on a fork of pi-mono. Within a week of launch, it was covered by multiple media outlets.

Official results: Opus 5 + Prime Agent scored 95.5% on ARC-AGI-3, surpassing the human expert baseline of 95.4%. In our field test, it wrote working Rust emulators for the SEGA Genesis and Game Boy Color without reference code.


Two Core Abstractions

Prime Agent's design revolves around two paper-grade abstractions, which is also its fundamental divergence from traditional agent harnesses (fixed tool schemas + context compression).

1. Recursive Language Model (RLM)

Context is a variable, and a subagent is a function call.

A traditional harness gives the model a pile of fixed-schema tools (read_file, run_command...) and then uses context compression to "help" the model save tokens. RLM does the opposite:

  • The model has only one tool: a persistent IPython kernel
  • Context is handled as a variable (prompt-as-a-variable); message history is stored in variables, and the model can write code against its own context
  • A subagent is just a function call: rlm("subtask") spawns a complete subagent (with its own model, kernel, and history) and returns non-blockingly, with the result delivered asynchronously

The effect: in theory it supports infinitely long sessions without losing history, because history is not "compressed away" but "referenced programmatically".

2. Continual Harness (the Self-Evolving Shell)

The harness state H = (prompts, sub-agents, skills, memory), and the Agent can CRUD it.

The /refine command reads the agent's own execution traces, extracts lessons, and makes minimal, evidence-backed updates to the harness state:

  • Update types: supplementing prompts, memory, skill descriptions, subagent specs
  • The base system prompt is never mutable (the safety anchor)
  • Every update records a snapshot and can be rolled back by ID
  • Planning runs in the background without blocking the conversation

This is the engineering realization of "an agent improving itself": not changing weights, but changing its own operating environment.


Other Engineering Features

FeatureDescription
Daemon-backed sessionsThe agent keeps running after the terminal disconnects, and you can reattach at any time
Direct inter-agent communicationRunning agents can discover each other, exchange messages, and work together
Heartbeats & SchedulesPeriodic/scheduled re-entry into a session
Persistent Goals/goal keeps a goal alive across turns until it is complete
Bounded Autonomous Mode/autonomous executes autonomously within a turn/token/time budget + a custom quality gate
Skills = executable Python packagesA built-in skill creator distills workflows into skills
Headless modeJSON / RPC modes for automation integration

🪞 Comparison with UltraClaw Infrastructure (the Focus of This Article)

On our own OpenClaw system, we assembled nearly identical capabilities using external infrastructure. Here is the comparison table:

Prime Agent (harness-native)UltraClaw (OpenClaw + plugin layer)
Continual Harness /refineThe three-piece suite of agent-evolver + skill-curator + SkillOpt
RLM subagent function callssessions_spawn + subagents orchestration
Daemon + reattachThe OpenClaw Gateway resident process + sessions
Heartbeats / SchedulesOpenClaw cron + heartbeat (the same mechanisms)
Skills = Python packagesThe SKILL.md skill system + SkillOpt nightly optimization
Snapshot + rollbackOur verify.py + Gate JSON validation
Inter-agent communicationsessions_send + multi-agent routing

Conclusion: the direction is exactly the same, which proves we are on the frontier path. The difference is at the implementation level:

  1. Prime Agent builds self-evolution into the harness kernel (/refine is a first-class citizen); ours is a plugin layer stacked on top (three independent skills + external cron)
  2. Prime Agent's RLM solves long-task memory with "programmatic context"; we use a three-layer file system of vector memory + daily log + session transcript
  3. Our advantages: multi-channel integration (Feishu/WeChat/WhatsApp), a business-scenario skill library (400+ skills such as AK-SDD Hong Kong stock research), and a compliance audit layer. These are vertical accumulations a general-purpose harness will not build

One interesting observation: Prime Agent's /refine never touches the base system prompt, and the pit we fell into (Lesson #76-77: model switching causing behavioral regression) precisely demonstrates the importance of an immutable anchor. Different teams independently converging on the same design suggests this is the correct answer.


🔬 Field Test Record (Mac Studio, early hours of 2026-08-13)

Installation Audit

The official install script is 1,620 lines; we audited it before installing (rather than blindly running curl | sh):

  • SHA-256 verification of the release package ✅
  • Sudo is needed only when Node.js is missing ✅
  • No suspicious outbound connections or persistence behavior ✅

Installation result: prime-agent v0.7.2, 192 dependency packages, checksum verification passed.

Connecting to Our Own API Relay Layer (Sub2API)

Prime Agent supports adding any OpenAI-compatible provider via ~/.prime/agent/models.json:

{
  "providers": {
    "sub2api": {
      "baseUrl": "http://localhost:8080/v1",
      "api": "openai-completions",
      "apiKey": "sk-...",
      "compat": {
        "supportsDeveloperRole": false,
        "supportsReasoningEffort": false
      },
      "models": [
        { "id": "deepseek-v4-pro", "reasoning": false, "contextWindow": 128000 }
      ]
    }
  }
}

Pitfall: Cloudflare Error 1010

The first test returned 403 Your request was blocked. The troubleshooting process:

  1. curl direct to the API → 200 ✅ (ruling out a key issue)
  2. curl simulating stream/tools/UA → all 200 ✅
  3. A local capture server grabbed prime-agent's complete request → replayed it to the production domain → 403, error code: 1010

Root cause: Cloudflare WAF's rule 1010 (blocking by browser signature). Prime Agent uses the OpenAI JS SDK, with a User-Agent of OpenAI/JS 6.47.0 plus a string of X-Stainless-* fingerprint headers, which the WAF identified as bot SDK traffic and blocked.

Solution: switch to a local endpoint (http://localhost:8080/v1) to bypass Cloudflare. Sub2API runs on the same machine anyway, so latency is lower.

Field Test Results

Test 1 (basic connectivity):

> Reply with exactly: PRIME_AGENT_TEST_OK
< PRIME_AGENT_TEST_OK  ✅

Test 2 (a real coding task): asked to create a memoized version of fibonacci and verify fib(50):

  • The Agent wrote files, executed, and verified through the IPython kernel
  • It produced fibonacci.py (an @lru_cache implementation, with boundary checks and a docstring)
  • It correctly returned fib(50) = 12586269025 ✅

The whole process took about 40 seconds, passed on the first try, and needed no human intervention.


⚠️ Risks and Notes

  1. It is not a security sandbox: the official README explicitly warns that it executes model-generated Python with the user's privileges. The worker/kernel processes have lifecycle isolation, but that does not constitute a security boundary. Use an external sandbox to run untrusted code.
  2. Cost: it is designed for frontier models, and the token consumption of long tasks + recursive subagents is considerable.
  3. A young ecosystem: it was open-sourced only on August 6; the documentation is complete, but community examples are still few.

What It Means for Us

  1. Borrow /refine's snapshot + rollback mechanism: our SkillOpt already has a gate, but lacks the fine-grained control of "rolling back a single update by ID"
  2. RLM's programmatic context is worth tracking: when OpenClaw supports it, long-task memory may no longer need a three-layer file system
  3. Cloudflare WAF will falsely flag legitimate agent traffic: this is a new pitfall at the infrastructure layer, worth adding to the operations manual
  4. Self-improvement is not the future, it is now: Prime Agent, our seven-piece suite, and SkillOpt are all converging on this direction in 2026

Prime Agent is worth tracking continuously. Our next step: run a real long task in an isolated environment (for example, a complete AK-SDD research workflow) and test whether the Continual Harness's /refine can produce valuable self-improvement.


This article is based on a field test of Prime Agent v0.7.2, 2026-08-13.