OkHuman Security Review: 12 Defects Across 3 Severity Tiers
When an Agent framework can execute arbitrary bash commands on the local machine, every line of its code can be an attack surface. This article uses OkHuman as a case study to fully reproduce a security review methodology for open source Agent frameworks: from attack surface inventory, defect severity classification, evidence chain construction, to remediation recommendations. All findings include file and line numbers, so readers can verify them independently. This is not only a security report, but also a reproducible review process applicable to any team evaluating or deploying an Agent framework.
Why review an Agent framework
The security posture of Agent frameworks is completely different from that of traditional software. The attack surface of traditional software is usually a clear API boundary, inputs undergo schema validation, and outputs undergo template rendering. Agent frameworks, by contrast, allow models to drive command execution, read and write files, call external services, and even modify their own prompts. This means the consequences of "judgment errors" and "malicious input" are amplified by the same execution engine, and the amplification factor is difficult to predict.
OkHuman is a personal Agent runtime implemented from scratch using the Go standard library. It advocates "one process = one agent" and uses "a single bash meta-tool + HTTP-decoupled plugins" as its capability extension model. Its codebase is about 15,500 lines, has zero external dependencies, and binds to 127.0.0.1 by default. This design is positioned as "local single-user," but the purpose of a security review is precisely to test whether this default is sufficient.
We chose OkHuman as a case study for two reasons. First, its design is opinionated: zero dependencies, plugin decoupling, and atomic write persistence, all of which are engineering practices worth learning from. Second, its size is moderate: 15,500 lines are enough to cover the core patterns of an Agent framework, yet not so large that it cannot be reviewed line by line. For a project established only 7 days ago with only 9 stars, the level of completeness is unusually high, but high completeness does not equal high security.
Review Methodology: How to Systematically Review an Agent
We adopted a purely static, read-only review approach and did not modify any files in the repo. The review covered three dimensions, and each dimension had a clear checklist.
First, the command execution path. The Agent's core capability is executing commands, so we traced the complete chain line by line from the HTTP endpoint to bash execution, checking parameter validation, permission control, and output handling. Specifically, we traced how a request to the /chat endpoint passes through model decision-making and ultimately triggers the bash tool, as well as how the command text is written to a temporary file, how it is executed, and how the output is truncated.
Second, HTTP interface security. OkHuman exposes 18 endpoints. We checked the routes and handler one by one, examining authentication, authorization, Origin validation, and Content-Type validation. We used grep to search for keywords such as Authorization, Bearer, middleware, CORS, Origin, and Referer, to confirm whether any authentication mechanism exists.
Third, prompt injection and isolation. The Agent's conversation context receives user input, tool output, and external web page content. We checked whether these sources have isolation markers and whether the system prompt exposes sensitive information. We paid particular attention to the "life self-awareness" mechanism, namely whether the system prompt injects information such as the agent's own PID, port, path, and other details.
The review output is a structured defect list. Each defect includes: severity rating, evidence (file:line number), attack scenario, impact assessment, and remediation recommendation. Severity is divided into three levels: High risk (can cause silent failure or directly lead to the system being controlled), Medium risk (missing functionality or information leakage), and Low risk (engineering hygiene and maintainability issues).
Attack Surface Overview
OkHuman's attack surface can be divided into three layers: the externally reachable HTTP interface, model-driven command execution, and prompt injection in the conversation context.
The HTTP interface binds to 127.0.0.1:8451 by default, but the code has no enforced loopback mechanism. If the deployer changes server.host to 0.0.0.0, all endpoints are immediately exposed externally. More critically, all 18 endpoints are dispatched directly, without any authentication or authorization middleware, and without Origin or Host validation (internal/server/server.go:546-594). We searched with grep Authorization|Bearer|middleware|CORS|Origin|Referer server.go, and the result was zero hits, confirming this.
The command execution path enters through the /chat endpoint and, after model decision making, triggers the bash tool. The design in which command text does not enter argv is correct (internal/tools/tools.go:140-156), and it prevents pgrep -f self-matching; the separate process group and timeout SIGKILL are also solidly implemented (tools.go:158, tools.go:172-178). Output is capped at 32MB per stream; when the limit is exceeded, the pipe continues to be drained so the child process does not block (tools.go:162-163, tools.go:210-244). However, because the HTTP interface has zero authentication, any local webpage can use CSRF to drive the agent to execute arbitrary commands.
On the conversation context side, external content (tool output, background notifications, web scraping) enters without any isolation markers, leaving the prompt injection surface fully open. The system prompt even injects its own PID and port, expanding the blast radius of injection (internal/context/dynamic.go:44-58). Each LLM call prepends it to the front of the system prompt (internal/context/manager.go:331-338); if the LLM endpoint is a remote API, local path and PID information will be leaked to the model service provider.
Level-by-level breakdown: Complete list of 12 defects
The review ultimately yielded 12 defects, divided by severity into three levels: 4 high-risk, 4 medium-risk, and 4 low-risk.
High risk (4 items)
D-1: Silent degradation of scout semantic search. When the scout plugin starts llama-server, it does not specify --pooling, causing all embeddings to fail, but the system only logs one line and continues running in keyword mode. The health endpoint returns {"model":"on"}, misleading operators into thinking everything is normal. The root cause is that llama-server's default pooling is none, not OAI-compatible mode, so /v1/embeddings returns a 400 error.
D-2: Misleading health endpoint. The health endpoint conflates model: on with embedding: healthy; even if all embeddings fail, it still returns a normal status. This is a classic "silent failure" problem, and operators will mistakenly believe the system has been fixed.
D-3: Misjudgment of completion status. The test sleep 45 shows that after a task is moved to the background, the agent immediately replies "completed", but the task has not actually finished executing. The problem is not the mechanism (the mechanism that moves a foreground timeout to the background is itself correct), but the model's judgment: it mistakes "handed off to the background" for "already completed".
D-4: Zero authentication on the HTTP interface + CSRF can drive arbitrary bash. All 18 endpoints are dispatched directly, with no authentication middleware (internal/server/server.go:546-594). tryReadJSON does not check Content-Type (server.go:1591-1595), and any decodable body is treated as JSON. An attacker only needs to trick a user into visiting a malicious web page to drive the agent via CSRF to execute arbitrary commands. DNS rebinding also applies, because there is no Host validation.
Medium risk (4 items)
D-5: Absence of /proc on macOS causes life self-awareness to fail. The portToPID function scans /proc/net/tcp and all process fds, but macOS has no /proc, causing the function to always return -1 (internal/context/dynamic.go:66-100). The life-awareness block shows "unknown", and it only works on Linux.
D-6: GET /config returns api_key in plaintext. The configuration returned by the /config endpoint includes llm.api_key in plaintext (server.go:1403, server.go:221-223). The APIKey string in internal/llm/client.go:29 has no json ignore tag. Combined with D-4, an attacker can obtain the LLM API key with a single GET request, causing financial loss.
D-7: No web search plugin at all. grep web_search|brave|tavily|serp has zero hits across the entire repository. One of the most common agent needs is completely missing. When asked, the model directly answers "I do not have web search capability".
D-8: OkHuman misdiagnoses its own problems. Testing found that OkHuman mistakes the local scout plugin for a web search tool and draws incorrect conclusions based on this false premise. This is a diagnostic blind spot that misleads users.
Low risk (4 items)
D-9: The voice-chat plugin directory is empty. The README describes the functionality in detail, but the directory contains 0 Go files, so the documentation and implementation are inconsistent, which misleads users.
D-10: No CI automation. There are 9 test files but no GitHub Actions, so tests cannot run automatically. The existing tests are of good quality, but without automation they deliver no value.
D-11: Skewed test coverage. agent.go has 789 lines but only 172 lines of tests; server.go has 1,619 lines but only 334 lines. Core modules are where risk is concentrated, so test coverage should be filled in as a priority.
D-12: Hot reloading prompts has no version history. POST /prompts/reload can hot reload prompts, but there is no version history mechanism, so if changes break something it is hard to trace back.
The Three Most Dangerous Flaws: Attack Chain Analysis
From an attacker's perspective, the three most dangerous flaws are D-4, D-6, and D-3, which can form a complete attack chain.
D-4 allows any arbitrary web page to drive the local agent via CSRF, which is equivalent to RCE. The attack scenario is as follows: the user visits a malicious web page, and an embedded form on the page submits JSON to http://127.0.0.1:8451/chat. Because the body is JSON and no custom header is required, it is a CORS simple request; tryReadJSON does not check Content-Type, so a text/plain form body is also decoded successfully. The response cannot be read, but the request has already taken effect, so a drive-by triggers the agent to run arbitrary commands.
D-6, combined with D-4, allows an attacker to obtain the LLM API key. A single GET /config retrieves the API key for a third-party LLM service, which can be abused for fraudulent usage. If the api_key points to a billing endpoint, it can also cause financial loss.
Although D-3 is not a direct security vulnerability, it can cause state inconsistency in automated scenarios. It appears complete externally, but in reality it is not, which can mislead downstream systems that rely on the agent's reporting.
The common trait of these three flaws is "silence": CSRF attacks leave no trace; an api_key leak leaves only a request record in the logs; and misjudgment of completion status stems entirely from model behavior, making it difficult to detect externally. A silent defect is more dangerous than an obvious one.
Fix Recommendations for Developers
For the defects above, we propose the following fixes, ordered by priority.
Priority One: Minimal Authentication and api_key Redaction. Generate a random token at startup, and require the Authorization header on all non-/health endpoints; or at least validate Host/Origin against expected values. tryReadJSON should require Content-Type: application/json. At the same time, in the response to GET /config, redact api_key to "" or "***" (keeping only the length or the last 4 characters); add json:"-" to llm.ClientConfig.APIKey. If server.host != 127.0.0.1, a token should be mandatory rather than relying on configuration discipline.
Priority Two: Make scout fail loudly. Do not silently degrade when embedding fails. The health endpoint should distinguish model: on from embedding: healthy; or refuse to start when the startup self-check fails, forcing the problem to be seen. scout should explicitly specify --pooling last (Qwen3-Embedding officially recommends last-token pooling).
Priority Three: Distinguish "Submitted" from "Completed". The notification prefix for background tasks should clearly indicate status (for example, "Moved to background, not yet completed"); or directly educate the model in the tool definition's docstring: "Moving to background does not equal completion."
Priority Four: macOS Support or Clear Labeling. If macOS is supported, portToPID should use lsof -i :port or ps instead; if it is not supported, the README should indicate "Linux only".
Priority Five: Engineering Hygiene. Complete CI (GitHub Actions runs 9 tests), complete voice-chat or remove its README, increase test coverage for core modules, and add version tracking to prompt hot loading.
Security Positioning Assessment for an "Experimental Framework"
OkHuman positions itself as an "experimental framework." Does this mean security can be set aside for now? Our answer is no.
The value of an experimental framework lies in exploring new design patterns, not in avoiding engineering discipline. OkHuman's zero dependencies, plugin decoupling, and atomic write persistence are all designs worth promoting. However, the completeness of these designs should not become a reason to ignore security fundamentals.
"Local single-user" is a reasonable deployment positioning, but it is not a security guarantee. CSRF can cross browser boundaries, plaintext api_key can be read by other processes on the same machine, and prompt injection can be triggered by malicious web pages. These risks still exist in "local single-user" scenarios, only with lower severity.
Our security positioning assessment of OkHuman is: acceptable under the default configuration (bound to 127.0.0.1, local single-user), but it needs to complete three basic defenses: minimal authentication, api_key redaction, and failing loudly. These fixes have low engineering cost, but they can significantly reduce the attack surface.
For everyone who uses or develops Agent frameworks, our recommendation is: treat OkHuman as an automation tool with "clear steps, command-based verification, and external review," and do not trust its self-reported "completion." At the same time, prioritize fixing "silent failure" issues, because failing silently is more dangerous than failing visibly.
General Implications for Readers
The review results for OkHuman have implications for everyone who uses or develops Agent frameworks.
First, "local single-user" does not equal "secure." Even if it binds to 127.0.0.1 by default, CSRF can still cross browser boundaries. Minimal authentication is a necessary defense layer; it cannot rely on configuration discipline.
Second, "silent failure" is more dangerous than "visible failure." Framework designers should proactively guard against it: either fail loudly or provide clear health indicators.
Third, prompt injection is a structural problem for Agent. External content enters the conversation context without isolation markers, so framework designers should provide isolation mechanisms that let developers selectively mark suspicious input.
Fourth, test coverage and CI are the security baseline. Test coverage for core modules should be completed as a priority and run automatically through CI. Without automated tests, there is no value in running them.
Fifth, documentation and implementation must be consistent. The empty directory for voice-chat and the misleading response from the health endpoint can both mislead users and increase operational risk.
Sixth, determining completion status requires external review. Do not trust an agent's self-reported "completion"; design mechanisms that let external systems verify whether a task is truly complete.
Next Steps
This article uses OkHuman as a case study to demonstrate how to systematically audit an open-source Agent framework. Readers can refer to the following resources to learn more about security and architecture details:
- AI Security Red Team Practical Guide: Learn how to conduct adversarial testing on AI systems and establish a systematic security review process
- Is Your Data Safe? Privacy Protection in the AI Era: Explores data handling and privacy risks in Agent frameworks, including session log persistence to disk and permission control
- Enterprise Data Security Basics: Understand data classification, access control, and audit logs from an enterprise perspective
- OkHuman Architecture Deep Dive: Understand OkHuman's design tradeoffs from an architectural perspective, including zero dependencies, plugin decoupling, and atomic write persistence
- OkHuman Capability Evaluation and Plugin Development: Learn about OkHuman's functional boundaries and extension mechanisms, including how we developed a web search plugin for it
More in Evidence
- OkHuman Architecture Deep Dive: A Minimalist Agent Framework Written in Go
- OkHuman Capability Evaluation: An Agent That Misdiagnoses Itself
- How We Fixed OkHuman Semantic Search and Built It a Web Search Plugin
- A Reality Check on Decision Models: Why They Seem Miraculous Online but We Measured Only 54%: A Full Comparison of JEV / LAYA / KEV / CLM-8B and a Deployment Formula