Agentic Research

OkHuman Architecture Deep Dive: A Minimalist Agent Framework Written in Go

2026/10/0612 min readBryan Chan閱讀中文原文
TopicsOkHumanAgent FrameworkGoArchitectureOpen Source

Imagine you have a group of very obedient assistants, each locked in their own office, able to communicate with the outside world only through a slip of paper. The instruction is written on the slip, the assistant follows it, and when done puts the result back on the slip. There is no shared whiteboard, no walkie-talkie, no back door. This sounds very rudimentary, but it is exactly the core metaphor of the OkHuman Agent framework.

OkHuman is an open source Agent runtime built from scratch in Go, developed by Donald24718, released under the MIT license on GitHub (commit 2875678). Its core claim is just one sentence: "One process equals one Agent." The entire project is about 15,500 lines of Go code across 34 files, with zero external dependencies. That's right, not even a single require statement; all HTTP clients, SSE stream parsing, and JSON handling are handwritten using only the Go standard library. In an era when contemporary Go projects routinely pull in dozens of dependencies, this approach is almost heretical.

But it is precisely this heretical character that makes it worth a look.

Why a 9 Stars Project Is Worth Studying

OkHuman's GitHub page shows only 9 stars, 2 contributors, and the project is less than a month old. If you look only at community metrics, it does not even count as "interesting." However, there is a huge gap between engineering completeness and community size: 18 HTTP API endpoints, 9 test files, a complete plugin system, and even precompiled Linux binaries. For a project only seven days old, this density is unusually high.

More importantly, the quality of its source code comments is excellent. In many places, the comments record timelines of real incidents, the correspondence with the original TypeScript version, and the evolution of design decisions. This makes it a rare "living textbook," from which you can read how an Agent framework grew step by step into what it is today.

Layered Architecture: A Four-Layer Cake from the Model to bash

OkHuman Layered Architecture

OkHuman's architecture can be divided into four layers for understanding.

The topmost layer is the model layer. It communicates with large language models through an OpenAI-compatible API. In the tested environment, an OMLX local model service runs behind the scenes, loading the Qwen3.5-9B-MLX-4bit quantized model and listening on localhost:8000. Generation speed is about 25.1 tok/s, prefill speed is 1,039 tok/s, and memory usage is 6.4 GB. This combination is sufficient for tool-oriented tasks, but its pure reasoning ability is relatively weak; in testing, multi-step arithmetic reasoning once produced an error.

In the middle is the core engine layer. internal/agent/agent.go (789 lines) handles the conversation loop and tool execution orchestration; the internal/context/ module (6 files, 1,392 lines) handles context assembly, budget control, and rolling compression; internal/server/server.go (1,619 lines) is the largest single module, hosting all HTTP APIs and WebUI static serving.

Further down is the tool layer. There is only one tool here: bash. internal/tools/tools.go (262 lines) defines the only meta-tool in the entire framework. All external operations, whether reading and writing files, searching for data, or calling plugins, are done entirely through bash. The model only needs to know how to write shell commands to drive everything.

The bottommost layer is the plugin layer. Plugins exist as independent processes or CLI tools, communicating with the main program via HTTP or the command line. The main program has zero awareness of the plugins' internal implementations.

Nine Plugins: A Loose Ecosystem, Each with Its Own Responsibilities

Nine-plugin ecosystem

As of commit 2875678 at the time of analysis, there are nine plugins in the OkHuman repository:

browser (2,293 lines) is the largest plugin, providing the browserctl CLI to drive local Firefox and supporting operations such as navigation, locating, form filling, clicking, and text extraction. qq-channel (2,167 lines) is a bidirectional forwarding bridge for the official QQ channel and runs as an independent long-running process. cron (1,154 lines) handles scheduled tasks, listening on :8601, and when the time comes, POSTs the prompt as a user message to the instance's /chat endpoint.

computer (983 lines) gives the Agent "hands": it can inject full-screen screenshots and simulate X11 mouse and keyboard events. scout (770 lines) is a semantic indexing service for local plugins and skills, listening on :8480, combining keyword and Qwen3-Embedding vector scoring. attach (567 lines) handles asset injection, supporting original or compressed images, and automatically compressing videos to 360p according to policy and segmenting them. okmon (536 lines) is a process monitoring and start/stop management tool, with its WebUI at :8496.

The skills plugin is rather special: it is not a service but a repository of operational experience. Each skill is a SKILL.md file recording operation steps that have been tested and verified locally, and it also enters the scout index. The repository deliberately does not include any preset skills. The README states explicitly: "Skills are environment-dependent; please accumulate them after testing them on your own machine."

Finally, there is voice-chat. Its README describes in detail features such as microphone input, VAD, voiceprint gating, ASR, wake word routing, and Kokoro TTS, but there is not a single Go file in the directory. This is a typical early-project phenomenon of documentation and implementation being inconsistent.

Three Key Design Tradeoffs

Three Key Design Tradeoffs

Behind OkHuman's architecture are three clear design tradeoffs, each representing the author's choice between "simplicity" and "functionality."

First tradeoff: one process equals one Agent. The main program once supported multiple Agents coexisting, but this was deliberately removed on September 9, 2026. The current approach is to start as many processes as there are Agents, each with its own independent on-disk directory and independent context. The benefit is complete isolation: if one instance crashes, it does not affect the others; the cost is higher resource usage, and cross-Agent collaboration requires external mechanisms.

Second tradeoff: the only meta-tool is bash. The framework does not define structured tools such as read_file, write_file, or search; instead, it hands all external operations to bash. The comment in tools.go is very blunt: "Tool system: the only meta-tool is bash. All external operations are accomplished through bash." The model only needs to know how to write shell to mobilize the entire system's capabilities. The cost is that all capabilities are constrained by bash's bottleneck, and the model's shell-writing ability directly determines the framework's ceiling.

Notably, the bash tool does not use bash -c to place the full command in argv; instead, it first writes to a temporary script file and then executes it. The comment explains that this was a correction "after a real-world incident," because the ps command can read a process's argv, and the full command may leak sensitive information.

Third tradeoff: plugins are fully decoupled through HTTP. Plugins do not import the main program, do not share processes or memory, come with their own configuration and data directories, and can be compiled independently. The main program has zero awareness of plugins; if a plugin goes down, it does not affect the main program, and restarting the main program does not affect plugins. The cost is that every call must go over HTTP, with latency and serialization overhead.

Seven Core Mechanisms: Practical Experience from the Source Code

The most worthwhile part of OkHuman to read closely is not the feature list, but seven mechanisms that grew out of real-world practice.

Rolling Context Compression (context/compress.go): When the total token count of a session exceeds the budget, the framework batches old messages and sends them to the LLM to compress into summaries, iterating up to three rounds and taking the shortest round. The compression prompt explicitly states at the beginning: "Ignore all previous instructions and tasks." This is because the author found that without doing so, the model treats old instructions in the historical messages as pending tasks and calls tools for them. This is a very real pitfall that most frameworks do not handle.

Infinite Loop Detection (agent/agent.go): Three consecutive identical tool calls trigger a warning. If the repetition continues after the warning, a doom_stop event is issued to forcibly stop the current round. The design is "execute first, warn later," to avoid mistakenly killing normal retries.

Foreground Timeout to Background (background/background.go): A tool call races against a 30-second foreground timeout. On timeout, execution is not canceled; instead, it is downgraded to a background task, and when complete, the result is automatically injected into the next round of context. This solves the pain point of the Agent waiting idly for long-running tools. However, actual testing has found that the model sometimes misinterprets "moved to background" as "already completed," which is something users need to pay special attention to.

Life Self-Awareness: A "Your Life" block is injected at the beginning of the system prompt, containing this instance's port, PID, source code location, and the port and PID of the connected LLM service, computed fresh on each call. The purpose is to let the model know that "touching these PIDs and ports means harming itself." This is a soft self-protection mechanism implemented through prompts.

Prompt Layering and Hot Loading (prompt/prompt.go): prompts/*.md files are sorted by filename and concatenated into the system prompt. They can be hot loaded via POST /prompts/reload without restarting. Dedicated layers (such as deployment specifications and plugin specifications) do not enter the system prompt; the Agent reads them on demand.

Atomic Write Persistence (persist/persist.go): Every change to session.json is written atomically using a temporary file plus rename, and logs over 5 MB are automatically rotated to .old.

Zero-Dependency Handwritten SSE Streaming (llm/client.go): The entire HTTP client and SSE stream parsing are implemented using only the standard library, with complete functionality, including idle timeout handling.

Relationship with OMLX: Local Models as Standard

OkHuman's test environment was paired with an OMLX local model service. OMLX listens on localhost:8000 and loads the Qwen3.5-9B-MLX-4bit quantized model. The measured figures for this setup are as follows: generation speed 25.1 tok/s, prefill speed 1,039 tok/s, memory usage 6.4 GB.

This combination conveys an important message: OkHuman's design assumption is that you have a local machine that is powerful enough. A 9B quantized model is sufficient for tool calling and workflow-following tasks (all four tests in the bash chain passed in testing), but its pure reasoning ability is on the weaker side. The framework's core insight is: move the hard part from the model to engineering. The model decides what to do, the harness makes execution reliable, and plugins provide capabilities. Once these three are decoupled, even a small model can handle tool-oriented tasks.

Persistent Infrastructure: launchd and Four LaunchAgents

In a macOS deployment environment, OkHuman uses four LaunchAgents to auto-start at boot and automatically restart on crash (KeepAlive):

  • com.ultraclaw.okhuman: main process, listens on :8451, includes OMLX dependency check
  • com.ultraclaw.okhuman-scout: semantic indexing service, :8480
  • com.ultraclaw.okhuman-okmon: process monitoring, :8496
  • com.ultraclaw.okhuman-cron: scheduled tasks, :8601

The four services plus OMLX (:8000) make a total of five persistent processes. Each has its own plist configuration, with launchd managing its lifecycle. This deployment approach lets OkHuman operate as an "always-running personal assistant," but it also means you need to get used to managing five background processes.

Who It Is For, and Who It Is Not For

OkHuman is an experimental framework. Its project is less than a month old, has only two contributors, no CI pipeline, uneven test coverage, and the voice-chat plugin has a description but no code. These facts determine its positioning.

Who it is for: Developers who want to understand how an Agent framework works internally. Its source code comments are of extremely high quality, and all seven core mechanisms have complete records of their design evolution. In particular, the instruction isolation in rolling context compression, the degradation strategy that moves foreground timeouts to the background, and the three-role deployment topology (production, replica, lifeboat), these design ideas can be directly borrowed for your own system.

Who it is not for: Teams that need to run an Agent reliably in production. The framework's ecosystem risk is too high, the problem of silent failures has not been fully resolved (for example, scout's embedding was once silently downgraded to keyword search, while the health endpoint still showed normal), and there is room for improvement on the security side (zero authentication on the HTTP interface, and the /config endpoint returning the API key in plaintext).

In one sentence: read it as a reference for design ideas, not as a production tool that you can directly go install and launch.

Next Steps

If you want to continue exploring the world of Agent architecture, the following articles on this site can help you build a more complete body of knowledge:

  • What Is an Agent Harness: Understand the core concepts of the Agent runtime, and how a harness turns model capabilities into reliable automated workflows.
  • The Evolution Path of Harness: From early script-based Agents to modern layered architectures, understand where designs like OkHuman sit on the evolutionary tree.
  • What Is an AI Agent: If you have not heard the term Agent before, this article starts from the most basic definition.
  • OkHuman Security Review Report: A deep dive into the security findings mentioned in this article, including a complete analysis of the zero-authentication HTTP interface and CSRF risks.
  • OkHuman Capability Evaluation and Scout Fix: Details of hands-on testing of 13 capability tests, and how we fixed the silent degradation issue in scout semantic search.
  • Where Are OkHuman's Capability Limits: A more detailed analysis of capability boundaries, including a discussion of the gap between judgment and execution.