How We Fixed OkHuman Semantic Search and Built It a Web Search Plugin
A 400 error, hidden beneath three layers of logs. The health endpoint said everything was normal, yet semantic search had already silently degraded to keyword matching. We spent an afternoon tracing this root cause chain, and in the end fixed it with a single line of code. Then we discovered: after the fix, this agent still could not search the web. So we built a search plugin from scratch, including three providers, an automatic fallback chain, and complete endpoint testing.
OkHuman is a personal Agent runtime implemented from scratch using the Go standard library. It advocates "one process equals one agent", the entire project has zero external dependencies, and its engineering completeness far exceeds its community size. This article records the complete process of the fix and plugin development, as well as the replicable methodology we took away.
Symptoms: health says it is fine, but it is actually broken
OkHuman's scout plugin handles two things: indexing local plugin and skill files, then providing semantic search capability. It uses llama-server to run a Qwen3-Embedding model, converting files into vectors, and calculating similarity scores during queries.
The problem is: semantic search never actually worked, but there was no alert.
scout's health endpoint returned {"model":"on"}. When operators saw this response, their intuitive reaction was "the embedding model is loaded, everything is normal". We initially judged it the same way. After all, the health endpoint is meant to confirm system status; if it says normal, then it should be normal, right?
But we decided to take one more step. We actually called /v1/embeddings:
POST /v1/embeddings
→ 400 Bad Request
→ Pooling type 'none' is not OAI compatible.
There was constant failure with embedding. For every indexing run and every query, embedding returned a 400 error. The way scout handled it was to log one line and then silently degrade to pure keyword matching. On the surface, search still worked; in reality, the semantic component never took part at any point.
This is the most dangerous part of the problem: not a feature crash, but a silent degradation that looks like everything is normal.
Root cause chain: a missing parameter
The root cause was one line. When starting llama-server, scout passed the --embedding parameter but did not specify the --pooling type. llama-server's default pooling is none, and this value is incompatible with the OpenAI-format /v1/embeddings endpoint.
// Before fix (plugins/scout/main.go approx lines 216-218)
"--embedding", "--no-webui",
// After fix
"--embedding", "--no-webui", "--pooling", "last",
Why choose last? Because Qwen3-Embedding officially recommends last-token pooling. We also tested mean, and it worked just as well, but last is the default recommended in the documentation.
There is a lesson worth remembering here: a health endpoint saying model:on only means the model files have been loaded onto the GPU. It does not mean the embedding endpoint is available. These are two different things. model:on only confirms that the model weights have been loaded into memory, but whether the embedding endpoint can properly return vectors still depends on whether the pooling configuration is correct. The only way to confirm that embedding actually works is to actually call /v1/embeddings and check the dimensions of the returned vector.
We ourselves almost made the same mistake as the framework: we looked only at a single cheap status metric and judged it a pass, without testing the actual functionality. Both frameworks and operators can make this mistake, so frameworks should actively guard against it, rather than shifting responsibility to operators' vigilance.
A One-Line Fix, Then Verification
After adding --pooling last, recompile scout, and verify end to end:
✅ POST /v1/embeddings → returns a real 1024-dimensional vector (non-degenerate)
✅ scout reindex → {"indexed":11,"model":"on"}
✅ semantic query → score 0.6139 (breakdown.sem=0.7697, the semantic component is indeed included in the calculation)
All three checks passed. Semantic search is back.
Note how we verified it: not by looking at the health endpoint, not by looking at the startup logs, but by directly calling the functional endpoint and checking the output values. This is the most central methodology of the entire article, and we will emphasize it again at the end.
Still Missing One Piece: agent Does Not Know It Can Search
After scout was fixed, we found a second problem.
OkHuman originally had no web search plugin at all. In the entire repository, grep web_search|brave|tavily|serp returned zero hits. When the agent was asked, "Help me search for a topic," it would directly answer, "I don't have web search capability."
We decided to build our own search plugin. But before building the plugin, there was a smaller thing to do first: we added five lines of search guidance to prompts/02-directories.md, clearly telling the agent where the search plugin is and how to use it, and added a red-text reminder: "Do not answer, 'I don't have web search capability.'"
This fixes the problem of the agent not knowing it has search capability. No matter how well the plugin is built, if the agent does not know to use it, it may as well not exist.
OkHuman's prompt system uses a layered design: prompts/*.md are concatenated in filename order into the system prompt; currently there are only two files: 01-identity.md defines agent identity, and 02-directories.md tells the agent where to find tools. We added search guidance to the second file, ensuring the agent can see this instruction in every turn of the conversation.
Endpoint Testing in Practice: No Assumptions, Just curl
The first step in building a custom search plugin is deciding which search provider to use. We chose three: Tavily (designed specifically for LLM agents), Brave (independent index), and DuckDuckGo (does not require an API key).
Tavily and Brave both require keys, which is straightforward. Tavily is a search API designed specifically for LLM agents, and its returned results have already been optimized for summarization. Brave has its own independent index and does not rely on Google or Bing. What really took time was the endpoint testing for DuckDuckGo. We did not assume any endpoint was available; instead, we ran curl on each one:
lite.duckduckgo.com/lite/?q= GET → ❌ Blocked (14KB almost entirely blocked)
html.duckduckgo.com/html/ POST → ✅ Works normally (result__a hit, 0 blocked)
api.duckduckgo.com JSON → ⚠️ Only instant answer, narrow coverage
lite version is blocked when using GET requests, and almost all content is filtered. The html version works completely normally when submitted via POST form. The api version returns JSON, but it only covers instant answer, and its search coverage is too narrow.
Ultimately, we chose to POST to html.duckduckgo.com/html/. There is another detail: links in the results are redirects in the /l/?uddg=<encoded> format, and you need to URL-decode the uddg parameter to get the real URL.
None of this can be learned from reading documentation. You only find out by actually making calls.
Three-Provider Plugin: Automatic Fallback Design
The plugin architecture follows OkHuman's decoupling contract: independent Go module, independent compilation, does not import the main program, does not share process or memory, and includes its own config.json for key management.
The three providers each have a different positioning:
Tavily is a search API designed specifically for LLM agents. Its returned results have already been summary-optimized, with a measured speed of 2.2 seconds. Brave has its own independent index and does not rely on Google or Bing. DuckDuckGo scrapes via an HTML endpoint, requires no key at all, and has a measured speed of 1.1 seconds, making it the fastest of the three.
The fallback chain is designed as follows: by default it uses Tavily. If Tavily fails (invalid key, timeout, rate limiting), it automatically tries Brave. If Brave also fails, it falls back to DuckDuckGo. Paid providers provide quality, and the keyless provider acts as the backstop.
Plugin usage:
search-bin "query string" # default tavily, 5 results
search-bin "query string" --n 10 # specify number of results
search-bin "query string" --provider ddgs # no API key required
search-bin "query string" --json # JSON output
search-bin --doctor # environment self-check
The deliverables include six files: search.go (356 lines, three providers plus automatic fallback plus normalized parameters), go.mod, config.json, README.md, meta.json, and the compiled search-bin (9.4 MB).
End-to-End Proof: Without Telling It the Path
After the plugin was built, the most critical test was: not telling the agent where the plugin was and letting it find and use it on its own.
We asked directly: "Search the web for 'Anthropic MCP' for me, and then tell me the titles of the first two results."
The agent located the search plugin on its own, called search-bin, and returned real search results after 13.5 seconds. The titles were correct, and the content was correct.
This test proved three things: the search plugin worked correctly; the fallback chain worked correctly; the five lines of guidance in the prompt were enough for the agent to autonomously find the plugin without knowing its path.
It is worth mentioning that OkHuman's tool model is a "single bash meta-tool", and all external operations are performed through bash. This means that after the agent finds the plugin, it calls the search-bin command line through bash to perform the search. The entire chain, from natural language understanding, to locating the plugin, to assembling the command line, to parsing the search results, is completed entirely by a 9B local model.
After reindex, scout's semantic index can also correctly index the new search plugin. The two fixes form a closed loop: scout can index the new plugin, the new plugin provides search capability, and the search results are in turn indexed through scout.
The Methodology to Take Away
The whole process can be distilled into four principles.
First, measure, do not guess. The health endpoint said model:on, and we almost believed it. The only thing trustworthy was actually calling the functional endpoint and checking the output values. Any declaration that "the status is normal," if it has not gone through functional-level verification, could be wrong.
Second, root cause chain. There may be several layers between the symptom and the root cause. The root cause of the 400 error was not "the embedding is broken," but "a parameter was missing at startup." Stopping at the first layer will make you fix the wrong thing. Every additional layer of asking "why" brings you one step closer to the real fix.
Third, minimal fix. The final fix was only one line of code. A good fix is not adding more code, but finding the smallest, precise point of change. Do not use ten lines for a problem that one line can solve.
Fourth, end-to-end verification. After fixing, do not just run a unit test and stop. From the user's perspective, walk through the entire flow and confirm that the final output is correct. The reason we did not tell the agent the plugin path was to verify the real end-to-end behavior, not to verify a shortcut that bypasses it.
These four principles apply not only to OkHuman. Any system involving multiple layers of abstraction will encounter the problem that "the upper layer thinks the lower layer is normal, but the lower layer has actually silently failed." The health endpoint is a layer of abstraction; it tells you "everything is normal," but this normal may only be partially normal. The embedding model loading is normal, but a pooling configuration error makes the endpoint unavailable. If you only check the upper layer, the problem will silently exist in the lower layer.
The solution is always the same: do not trust status declarations; measure actual behavior. Start verifying from the lowest-level functional endpoint, and confirm layer by layer upward, until the final user-visible output. Only then can you be sure that the system is truly normal, and not merely "looks normal."
Every step in this root cause chain can be reproduced independently: first start llama-server without --pooling, use curl to hit /v1/embeddings once and see it return 400, then add --pooling last and retry. The difference between the two responses is the most direct evidence.
Next Steps
This fix gave us a deeper understanding of OkHuman's plugin architecture and decoupling design. If you are interested in the architectural details behind it, the following related articles will help you understand the whole system more completely:
- OkHuman Architecture Deep Dive: From a zero-dependency Go implementation to the design trade-offs of one process per agent
- OkHuman Security Review Report: Zero HTTP authentication, CSRF risks, and a minimal authentication approach
- OkHuman Capability Evaluation: Complete data for 13 decidable tests
- MCP Protocol Guide: Another approach to plugin decoupling
- Function Calling Basics: A comparison from the sole bash meta-tool to a multi-tool model
- What the Agent Actually Sees: Understanding how prompt guidance affects agent behavior
More in Evidence
- OkHuman Architecture Deep Dive: A Minimalist Agent Framework Written in Go
- OkHuman Capability Evaluation: An Agent That Misdiagnoses Itself
- OkHuman Security Review: 12 Defects Across 3 Severity Tiers
- A Reality Check on Decision Models: Why They Seem Miraculous Online but We Measured Only 54%: A Full Comparison of JEV / LAYA / KEV / CLM-8B and a Deployment Formula