The Engineering of AI Memory Retrieval: A Full-Matrix Field Report on Four Access Paths × Ten Scenarios
Core question: How does an AI assistant that restarts every 24 hours reliably "remember everything"? Test environment: Mac Studio M3 Ultra · 96GB RAM · 1,193 lines of long-term memory · 18 in-depth Hong Kong stock research files · 7 access paths Methodology: Scenario matrix method, 10 real usage scenarios × 7 access paths = 70 cells of field data
Introduction: Memory Is Not One Problem, It Is Ten
People who build AI memory systems all make the same mistake at the start: assuming that "the memory problem" is one problem.
It is not. It is at least ten.
"When did the boss say to follow up?" is a task-tracking problem. "Why did we stop using Lovable to build the website?" is a historical-decision problem. "Who is Ray Leung, and what is our relationship with him?" is a person-association problem. "What are the red lines in the deployment process?" is a rule-retrieval problem. These queries differ completely in semantic structure, time span, and information density.
"No single path can do well in every scenario. The engineering challenge of a memory system is not choosing the best path, but routing each scenario to the right one."
This article breaks down the complete memory retrieval architecture of the Junze Zhiku AI assistant (UltraClaw) into seven access paths, runs a full-matrix comparison across ten real everyday work scenarios, and reveals the applicable boundary, fatal blind spot, and optimal combination for each path.
1. The Seven Memory Access Paths
In our architecture, the AI assistant can read memory in the following seven ways:
P1 · Direct File Read
The most primitive but most reliable method. The AI opens .md / .json / .txt files in the filesystem directly through the read tool.
| Dimension | Assessment |
|---|---|
| Latency | ~50ms |
| Data source | MEMORY.md (1,193 lines), daily/YYYY-MM-DD.md, follow_up_tracker.json, etc. |
| Core strength | Precise, complete, structured |
| Core weakness | You must know the file path; cannot associate across files |
P2 · Qdrant Vector Search
Calls a local Qdrant instance (the openclaw_mem collection) through vector-memory__mem_search, using the BGE-m3 model (1024 dimensions) for semantic similarity retrieval.
| Dimension | Assessment |
|---|---|
| Latency | ~100ms |
| Data source | 88 Markdown files → 1,156 vectors (continuously incrementally synced after a full import) |
| Index model | BGE-m3 (BAAI), 1024 dimensions, COSINE distance |
| Core strength | Cross-file semantic association; no need to know the precise location |
| Core weakness | Can only recall "similar content"; cannot reconstruct a structured schema |
P3 · MemoryHub Capture Pipeline
The capture_daemon scans Session JSONL + daily log files every minute, capturing conversation fragments and unstructured text into Qdrant (through its own deduplication and embedding pipeline). Keyword search is available via the memhub.best-thinktank.com/api/search API.
| Dimension | Assessment |
|---|---|
| Latency | ~500ms (API + network round trip) |
| Data source | Session JSONL (conversation records) + daily log (work journal) |
| Scan scope | ❌ Does not scan follow_up_tracker.json, pending_email_replies.json, memory/projects/ |
| Core strength | Automatic capture, no manual maintenance |
| Core weakness | Does not index structured task files; returns a lot of conversational noise |
P4/P5 · agentmemory Search
agentmemory is an external memory service based on a REST API (localhost:3111), offering two query paths: memory_recall (BM25) and memory_smart_search (hybrid search).
| Dimension | Assessment |
|---|---|
| Latency | ~200ms |
| Data source | The agentmemory internal index (~50 entries, manually imported) |
| Core weakness | Designed for the English environment of Claude Code, with a 0% hit rate for Chinese queries |
P6 · Session History
Reads the raw JSONL conversation records directly via sessions_history, precisely restoring the content of past conversations.
| Dimension | Assessment |
|---|---|
| Latency | ~300ms |
| Data source | agents/main/sessions/*.jsonl |
| Core strength | Precise conversation restoration |
| Core weakness | No structured summary; reading a large session is slow |
P7 · Bootstrap Injection
On startup, the AI automatically preloads the contents of SOUL.md, USER.md, MEMORY.md, RULES.md, PERMANENT-RULES.md, and similar files into the context window. Zero latency, zero API calls.
| Dimension | Assessment |
|---|---|
| Latency | 0ms |
| Preloaded file size | Configurable (currently an 80KB threshold) |
| Core strength | Automatically available in every conversation, zero cost |
| Core weakness | Truncated when files are too large; static content, requiring a restart to update |
2. The Ten Memory Retrieval Scenarios
These scenarios are extracted from UltraClaw's field work logs over the past 34 days, covering every type of memory need in an AI assistant's daily operation:
| # | Scenario | Trigger Frequency | Representative Query | Query Type |
|---|---|---|---|---|
| S1 | Startup load | Every session | "Who am I, who is the boss, what happened recently" | Identity + summary |
| S2 | To-do follow-up | Start of every conversation | "What needs to be reminded to the boss?" | Structured list |
| S3 | Project status | On demand | "How is the AIApps project going?" | Factual summary |
| S4 | Historical decision | On demand | "Why did we give up Lovable?" | Causal tracing |
| S5 | Pitfall lookup | Before similar tasks | "What pitfalls did we hit with DI HKEX data before?" | Lesson retrieval |
| S6 | Rules/process | While executing tasks | "What is the deployment process? What are the red lines?" | Precise rules |
| S7 | People/contacts | On demand | "Who is Ray Leung? What is our relationship with him?" | Entity association |
| S8 | Technical config | Troubleshooting/maintenance | "How is the cloudflared tunnel configured?" | Config lookup |
| S9 | Domain knowledge | Before research | "Which data center stocks have we researched before?" | Domain summarization |
| S10 | Cross-session continuity | Every conversation | "What were we talking about half-way through last time?" | Conversation restoration |
Scenario classification logic: S1-S2 are the "mandatory actions" of every conversation; S3-S9 are task-driven "on-demand queries"; S10 is the "continuity need" of the time dimension. These three categories place completely different demands on a memory system.
3. Full-Matrix Field Results
Below is a field comparison of hit rate, precision, and completeness for the seven paths across the ten scenarios. The test queries are all taken from real conversation scenarios that actually occurred.
3.1 S1 · Startup Load
Query: "Who am I, who is the boss, what happened recently"
| Path | Hit Rate | Precision | Completeness | Latency | Assessment |
|---|---|---|---|---|---|
| P7 Bootstrap | 100% | 🟢🟢🟢 | 🟢🟢🟡 | 0ms | Primary, SOUL/USER/MEMORY/RULES injected automatically |
| P1 File | 100% | 🟢🟢🟢 | 🟢🟢🟢 | ~50ms | The underlying source of Bootstrap |
| P2 Qdrant | 80% | 🟢🟢🟡 | 🟡 | ~100ms | Can fill the gap after MEMORY.md truncation |
| P3 MemHub | 30% | 🟡 | 🔴 | ~500ms | Conversational fragments, no structured identity information |
| P4/P5 agentmemory | 0% | 🔴 | 🔴 | ~200ms | Ineffective for Chinese queries |
🔬 Insight: Bootstrap injection is the lifeline of S1. MEMORY.md is currently 1,193 lines and still growing. Even though the threshold has been raised from 12KB to 80KB, there is still a risk of hitting the ceiling. Qdrant vector search is the best fallback: when the Bootstrap content is truncated, semantic search can restore the missing part.
3.2 S2 · To-Do Follow-Up
Query: "What needs to be reminded to the boss?" (the test file contains 12 tech items + 6 business items + 2 email to-dos)
| Path | Hit Rate | Precision | Completeness | Latency | Assessment |
|---|---|---|---|---|---|
| P1 File | 100% | 🟢🟢🟢 | 🟢🟢🟢 | ~50ms | Gold Standard, follow_up_tracker.json structured JSON |
| P2 Qdrant | 20% | 🟢🟢🟡 | 🔴 | ~100ms | Top 2 hits the core instruction, but misses all 18 specific tasks |
| P3 MemHub | 5% | 🟡 | 🔴🔴 | ~500ms | Only picks up the conversational fragment "the key reminder set up yesterday..." with no task list |
| P4/P5 agentmemory | 0% | 🔴 | 🔴 | ~200ms | Completely off-topic (returns unrelated stock report fragments) |
| P6 Session | 10% | 🟡 | 🟡 | ~300ms | Can see recent conversation but has no structured summary |
🔬 Insight: This is the classic showdown of structured vs unstructured.
follow_up_tracker.jsonis a carefully designed JSON schema containing precise task descriptions, priorities, and statuses for T1-T14. Vector search can "sense" the existence of tasks from the semantic space (a 20% hit rate), but can never reconstruct this schema. This is not a technical limitation; it is a fundamental difference in information structure. A numeric schema does not live in semantic space.
3.3 S3 · Project Status
Query: "How is the AIApps project going?"
| Path | Hit Rate | Precision | Completeness | Latency | Assessment |
|---|---|---|---|---|---|
| P1 File | 100% | 🟢🟢🟢 | 🟢🟢🟢 | ~50ms | memory/projects/AIApps.md + MEMORY.md #26 |
| P2 Qdrant | 85% | 🟢🟢🟡 | 🟢🟡 | ~100ms | Semantic search for "AIApps Flutter APK" works, but lacks the latest progress |
| P7 Bootstrap | 60% | 🟢🟡 | 🟡 | 0ms | A summary exists in MEMORY.md |
| P3 MemHub | 40% | 🟡 | 🟡 | ~500ms | May pick up related discussion fragments |
🔬 Insight: The best source for project status is
memory/projects/PROJECT.md(independently maintained) + MEMORY.md (the summary in Key Decisions). Qdrant can find related passages, but they may not be the most recent. When project progress is maintained in a structured way, files are the best path; when project discussion is scattered across conversations, Qdrant + MemHub have incremental value.
3.4 S4 · Historical Decision
Query: "Why did we give up Lovable?"
| Path | Hit Rate | Precision | Completeness | Latency | Assessment |
|---|---|---|---|---|---|
| P2 Qdrant | 100% | 🟢🟢🟢 | 🟢🟢🟡 | ~100ms | Best, semantic search for "Lovable data loss" perfectly hits MEMORY.md Key Decision -1 |
| P1 File | 70% | 🟢🟢🟢 | 🟢🟢🟢 | Variable | Requires first knowing that the answer is in MEMORY.md Key Decision -1 (indexed as -1!) |
| P3 MemHub | 30% | 🟢🟡 | 🟡 | ~500ms | Picks up fragments if it was ever discussed |
🔬 Insight: This is Qdrant's sweet spot scenario. Historical decisions are usually scattered across long files (MEMORY.md, 1,193 lines), and you do not know the precise location, or even whether it exists. But the keyword "Lovable" plus the semantic signature of a decision lets vector search hit in one shot. Direct file reading in this scenario instead requires manual localization: you must first know which file the answer is in, then search within 1,193 lines.
3.5 S5 · Pitfall Lookup
Query: "What pitfalls did we hit with AK-SDD DI HKEX data before?"
| Path | Hit Rate | Precision | Completeness | Latency | Assessment |
|---|---|---|---|---|---|
| P1 File | 100% | 🟢🟢🟢 | 🟢🟢🟢 | ~50ms | memory/lessons/index.md + the specific lesson files |
| P2 Qdrant | 75% | 🟢🟢🟡 | 🟢🟡 | ~100ms | Semantic search works, but may miss related pitfalls with similar numbering |
| P7 Bootstrap | 50% | 🟢🟡 | 🟡 | 0ms | The MEMORY.md Lessons section is in Bootstrap |
| P3 MemHub | 30% | 🟡 | 🟡 | ~500ms | Individual pitfalls mentioned in conversation |
🔬 Insight: Pitfall lookups have two subtypes: "discovery" and "confirmation". Discovery ("what pitfalls are there before doing DI?") suits Qdrant: you do not know what pitfalls exist, and semantic search helps you discover them. Confirmation ("what exactly caused the 9982 suspension misjudgment?") suits direct file reading: you know what you are looking for and need the precise content.
3.6 S6 · Rules/Process
Query: "What is the deployment process? What are the red lines?"
| Path | Hit Rate | Precision | Completeness | Latency | Assessment |
|---|---|---|---|---|---|
| P7 Bootstrap | 100% | 🟢🟢🟢 | 🟢🟢🟢 | 0ms | RULES.md + PERMANENT-RULES.md already injected |
| P1 File | 100% | 🟢🟢🟢 | 🟢🟢🟢 | ~50ms | skills/deploy-vercel/SKILL.md |
| P2 Qdrant | 40% | 🟢🟡 | 🟡 | ~100ms | Can find related passages but not the complete rule |
🔬 Insight: Bootstrap + skill files are the perfect solution for rule queries. Vector search performs notably worse in this scenario than in others, because rules must be precise and complete and cannot rely on "roughly similar". One rule missing one step can be a deployment disaster.
3.7 S7 · People/Contacts
Query: "Who is Ray Leung? What is our relationship with him?"
| Path | Hit Rate | Precision | Completeness | Latency | Assessment |
|---|---|---|---|---|---|
| P2 Qdrant | 95% | 🟢🟢🟢 | 🟢🟢🟡 | ~100ms | "Ray Leung Matrix Group HKOW cultural creative" hits precisely |
| P1 File | 60% | 🟢🟢🟢 | 🟢🟢🟢 | Variable | Scattered across MEMORY.md + daily/05-09.md and other files |
| P7 Bootstrap | 40% | 🟢🟡 | 🟡 | 0ms | Present if it is in MEMORY.md |
| P3 MemHub | 50% | 🟢🟡 | 🟡 | ~500ms | Fragments if it was mentioned in conversation |
🔬 Insight: People information is the scenario best suited to vector search: you do not know which file this name appears in (MEMORY.md? the 5/9 daily log? the HKOW project file?), but semantic search can associate across files automatically. This is Qdrant's second sweet spot.
3.8 S8 · Technical Config
Query: "How is the cloudflared tunnel configured?"
| Path | Hit Rate | Precision | Completeness | Latency | Assessment |
|---|---|---|---|---|---|
| P1 File | 100% | 🟢🟢🟢 | 🟢🟢🟢 | ~50ms | MEMORY.md Key Decision #27 |
| P2 Qdrant | 90% | 🟢🟢🟢 | 🟢🟡 | ~100ms | "Cloudflare Tunnel DNS LaunchAgent" hits precisely |
| P3 MemHub | 30% | 🟡 | 🟡 | ~500ms | If the configuration process was discussed |
🔬 Insight: In technical config scenarios, direct file reading and Qdrant perform comparably. Qdrant's unique value lies in cross-file association: the same query can hit the tunnel configuration + the DNS lesson + the LaunchAgent configuration at once, forming a richer context than a single file.
3.9 S9 · Domain Knowledge
Query: "Which data center stocks have we researched before?"
| Path | Hit Rate | Precision | Completeness | Latency | Assessment |
|---|---|---|---|---|---|
| P2 Qdrant | 90% | 🟢🟢🟢 | 🟢🟢🟡 | ~100ms | "1686 Sunevision 9698 GDS data center" hits semantically |
| P1 File | 70% | 🟢🟢🟢 | 🟢🟢🟢 | ~50ms | MEMORY.md Key Decision #50 |
| P7 Bootstrap | 50% | 🟢🟡 | 🟡 | 0ms | The research-target summary passage in MEMORY.md |
🔬 Insight: Domain knowledge is Qdrant's third sweet spot. The query "data center" and "1686.HK Sunevision" in memory are highly semantically related, even if the query term and the original text are not exactly the same. The obstacle to direct file reading is needing to know the answer is in Key Decision #50 and to understand the classification logic.
3.10 S10 · Cross-Session Continuity
Query: "What were we talking about half-way through last time?"
| Path | Hit Rate | Precision | Completeness | Latency | Assessment |
|---|---|---|---|---|---|
| P6 Session | 100% | 🟢🟢🟢 | 🟢🟢🟢 | ~300ms | sessions_history precisely restores the conversation |
| P3 MemHub | 60% | 🟢🟡 | 🟢🟡 | ~500ms | The capture daemon is designed specifically for this |
| P2 Qdrant | 30% | 🟡 | 🔴 | ~100ms | Present in conversation but not structured |
| P1 File | 10% | 🔴 | 🔴 | N/A | The daily log has a summary but it is not real-time |
🔬 Insight: Session history is the most authoritative cross-session source. The original purpose of the MemHub daemon was to fill this gap, but its current capture granularity is too coarse: it records "fragments" but not "context", and cannot answer questions that need a conversational flow such as "where did we get to".
4. The Composite Scoring Matrix
| Path \ Scenario | S1 | S2 | S3 | S4 | S5 | S6 | S7 | S8 | S9 | S10 | Total |
|---|---|---|---|---|---|---|---|---|---|---|---|
| P1 File | 10 | 10 | 10 | 7 | 10 | 10 | 6 | 10 | 7 | 1 | 81 |
| P2 Qdrant | 8 | 2 | 8 | 10 | 8 | 4 | 10 | 9 | 9 | 3 | 71 |
| P7 Bootstrap | 10 | 1 | 6 | 4 | 5 | 10 | 4 | 3 | 5 | 0 | 48 |
| P3 MemHub | 3 | 1 | 4 | 3 | 3 | 1 | 5 | 3 | 3 | 6 | 32 |
| P6 Session | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 10 | 11 |
| P4 recall | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| P5 smart | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
Scoring: 10 = perfect, 8-9 = excellent, 5-7 = usable, 3-4 = barely, 1-2 = insufficient, 0 = ineffective
5. Five Core Insights
Insight 1: Direct File Reading + Qdrant Is the Golden Combination
Direct file reading is unbeatable when "I know what to look up" (structured scenarios S2/S6); Qdrant is unbeatable when "I do not know which file the answer is in" (discovery scenarios S4/S7/S9). Together they cover the high-quality needs of 9/10 scenarios. They are not competitors but complements, like a reference book and a search engine: one gives you a precise answer, the other tells you where to find it.
Insight 2: The Gap Between Structured and Unstructured Cannot Be Crossed with Vectors
This is the most profound finding of this test. The T1-T14 task list in follow_up_tracker.json is a hand-designed JSON schema. Vector search can "sense" from semantic space that "there are to-dos", but can never restore the independence and priority relationships among T1 (ES container OOM fix), T2 (PG vector stringification bug), and T3 (daemon startup blocking).
"Numeric attributes are a product of the schema, not a product of semantics."
A memory system that relies only on vector search will inevitably fail in structured query scenarios (S2 to-do follow-up, S6 rule retrieval). Vector memory is necessary, but not sufficient.
Insight 3: The MemoryHub Capture Pipeline's Design Has Drifted from the Core Need
The core design of the MemoryHub capture daemon is "passively record conversation fragments" → "search the full text afterward". But in the field, the memory capabilities an AI assistant actually needs are:
- Structured task state: the daemon does not scan follow_up_tracker.json
- Cross-file relationship association: the daemon does not build an entity graph
- Real-time context queries: the API latency is 500ms+, too high for startup loading
It is currently more of a "memory black box" than a "memory tool". This is not to say it has no value: it provides a certain degree of cross-conversation visibility in S10 (cross-session continuity). But there is a significant gap between its design assumption ("passive capture is enough") and the actual need ("active indexing + structured understanding").
Insight 4: agentmemory Should Be Demoted from the Startup Flow
Across two independent comparison tests (20 queries in total), agentmemory's hit rate for Chinese queries was 0%.
The root cause lies in its design assumption: agentmemory was designed for English coding agents such as Claude Code, automatically capturing sessions through hooks and accumulating memory naturally. But our memory corpus is Traditional Chinese, structurally mixed, and domain-specific (Hong Kong stocks, M&A, Hong Kong law). agentmemory's embedding model and search algorithm were never tuned to handle this kind of data.
Recommendation: keep agentmemory but remove the mandatory check from HEARTBEAT.md and the startup flow. It is not a bad tool; it is just not suited to our scenario.
Insight 5: Bootstrap Injection Is the Invisible MVP
The automatic injection of SOUL.md + USER.md + MEMORY.md + RULES.md + PERMANENT-RULES.md provides completely zero-latency memory availability in S1 (startup load), S6 (rules and process), and part of S3 (project summary).
This is the highest-ROI mechanism in the entire architecture: zero query cost, zero maintenance cost, automatically available in every conversation. But it has a ceiling: file growth triggers truncation. The solution is not to raise the threshold without limit (a context window is not free), but to let the Qdrant vector fallback cover the truncated part.
6. Scenario Routing: Let Every Query Take the Right Path
Based on the insights above, here is a three-layer routing strategy for memory retrieval:
Incoming query
│
├─ Structured query (S2 to-do, S6 rules)
│ → P1 direct file read (follow_up_tracker.json / RULES.md)
│
├─ Discovery queries (S4 decisions, S7 people, S9 domain knowledge)
│ → P2 Qdrant vector search + P1 document confirmation
│
└─ Conversation Continuity (S10)
→ P6 Session History (Short) + P3 MemHub (Medium) + P1 Daily Log (Long)
Routing logic: use files for structured queries, vectors for discovery, and the timeline for continuity. Do not use vectors to search a task list, and do not use files to search cross-file associations.
7. Optimization Roadmap
🥇 P0: Immediate Impact (smallest code change, largest effect)
| # | Recommendation | Affected Scenarios | Expected Gain |
|---|---|---|---|
| 1 | Default every query to dual-track: direct file read + Qdrant vector search | S2-S5, S7-S9 | Recall +40% |
| 2 | Expand the MemoryHub daemon's scan scope: add follow_up_tracker.json, pending_email_replies.json, memory/projects/*.md | S2, S3 | MemHub S2 hit rate 5%→60% |
| 3 | Demote agentmemory from the startup flow: keep it but do not make it a mandatory check | Global | Reduce noise and hidden maintenance cost |
🥈 P1: Mid-Term Refactor (architecture-level improvements)
| # | Recommendation | Affected Scenarios | Expected Gain |
|---|---|---|---|
| 4 | Build a layered memory index: a structured layer (JSON task files) + an unstructured layer (vectors) + a temporal layer (Session JSONL), with automatic routing at query time | S1-S10 | Global recall +50% |
| 5 | Automatically tag tasks during conversation: when the boss says "remember this" → the daemon automatically extracts and writes it into follow_up_tracker | S2 | Automation rate 0%→80% |
| 6 | Incrementally index task files into Qdrant: every update to follow_up_tracker.json → automatic mem_save | S2 | Qdrant S2 hit rate 20%→70% |
🥉 P2: Long-Term Vision (paradigm upgrade)
| # | Recommendation | Affected Scenarios | Expected Gain |
|---|---|---|---|
| 7 | Upgrade MemoryHub from "passive recorder" to "active assistant": periodically scan all structured files and push summaries | S2, S3, S10 | From black box to tool |
| 8 | Cross-session intent tracking: automatically detect "the topic we left half-way through last time" and surface it at startup | S10 | Continuity from manual to automatic |
| 9 | A personal knowledge graph: automatically extract a person/company/project entity-relationship graph from MEMORY.md + daily log + projects | S4, S7, S9 | From keyword search to relationship discovery |
📈 Quantified Expected Gains
S2 To-Do Follow-Up (Current State → P0 → P1):
File 100% → 100% → 100%
Qdrant 20% → 70% → 90%
MemHub 5% → 60% → 90%
S4 Historical Decisions (Current State → P0 → P1):
File 70% → 70% → 80%
Qdrant 90% → 95% → 95%
MemHub 30% → 60% → 70%
S10 Continuity (Current State → P1 → P2):
Session 100% → 100% → 100%
Auto push 0% → 30% → 80% ← from 0 to 1
Conclusion: Memory Is Not a Database, It Is a Routing System
The core conclusion of this full-matrix test can be summed up in one sentence:
"A good memory system does not choose one best database to store everything. It routes each type of memory need to the path that suits it best."
Direct file reading suits precise structured queries. Qdrant vector search suits cross-file semantic discovery. Bootstrap injection suits zero-latency identity and rules. Session history suits conversational continuity. MemoryHub suits automatic capture.
Five paths, five roles. When compared as standalone solutions, each has blind spots. When treated as nodes in a routing system, each covers the shortcomings of another.
This is not a question of technical choice. It is a question of architectural design.
Methodology disclosure: All test data in this article comes from field testing by the Junze Zhiku AI assistant (UltraClaw) on May 23, 2026. The test queries are all based on real scenarios from the past 34 days of work logs. The scoring is a mix of subjective and objective: hit rate and latency are measured data, while precision and completeness are human assessments.
Test limitations: The path tests for agentmemory were limited by its indexing of only about 50 manually imported memories (vs Qdrant's 1,156). If agentmemory were given an equivalent full import, the results might differ, but given its original design intent (an English coding agent context), we believe that even with more data, the structural problems of Chinese semantic search would not fundamentally change.
More in Evidence
- A Reality Check on Decision Models: Why They Seem Miraculous Online but We Measured Only 54%: A Full Comparison of JEV / LAYA / KEV / CLM-8B and a Deployment Formula
- The "Non-Text-Generating Model": Jev and the New System One Category, and How Agent Architecture Changes When AI Only Answers Multiple Choice
- WeChat Open Source WeMM-Embedding Deep Dive: The Multimodal Embedding Model Topping MMEB-v2, Can It Run on Your Mac?
- A Source-Level Architectural Dissection of DeepSeek Harness: How an Everything-Is-a-Plugin Agent Framework Is Built