Agentic Research

We write down what we actually ran.

Architecture, memory, retrieval, models and toolchains for AI agents. Every piece states its method, raw data and sources, and carries a last-updated date.

Latest

A Reality Check on Decision Models: Why They Seem Miraculous Online but We Measured Only 54%: A Full Comparison of JEV / LAYA / KEV / CLM-8B and a Deployment Formula

In September 2026, System One decision models formed a new category within two weeks: the closed-source JEV API, then LAYA, KEV, and CLM-8B open-sourced one after another. This article does more than summarize the differences among the four; it uses six everyday scenarios to explain how they are actually used, and places the official marketing side by side with our same-question measurements: on the same set of questions, full-coverage accuracy was only 54%, but with confidence gating it reached 91.7%.

2026-09-27 · Evidence

The "Non-Text-Generating Model": Jev and the New System One Category, and How Agent Architecture Changes When AI Only Answers Multiple Choice

TypeSafe AI emerges from stealth with a $40M seed and releases Jev, a System One model that generates no text and only answers typed questions (choice / score / noul). It answers multiple questions in parallel in a single forward pass, returns calibrated probabilities, with ~100ms latency, priced at $0.042/MTok.

2026-09-19 · Evidence

WeChat Open Source WeMM-Embedding Deep Dive: The Multimodal Embedding Model Topping MMEB-v2, Can It Run on Your Mac?

Breaking down WeMM-Embedding from the arXiv 2608.24053 technical report and source code: Qwen3.5 native multimodal backbone, <embedding> token pooling, two-stage training (Semantic-ID resampling + bidirectional KL distillation + model merging), full MMEB-v2/v3 leaderboard results, and local Mac Studio MPS benchmarks (47ms text / 321ms image / 6

2026-08-31 · Evidence

A Source-Level Architectural Dissection of DeepSeek Harness: How an Everything-Is-a-Plugin Agent Framework Is Built

Dissecting the complete architecture of DeepSeek Harness (dsh) from 219 workspace packages and 450k lines of TypeScript: the Cordis plugin foundation, Event Sourcing session log, the three roles of capability seams, a single home for policy + fail-closed approvals, four layers of context engineering defense, and an equally plugin-based frontend. Includes a porting assessment for in-house Agent infrastructure.

2026-08-16 · Evidence

DeepSeek Harness Deep Research: Day-One Field Test of DeepSeek's Official "Everything Is a Plugin" Agent Framework

On August 13, 2026, DeepSeek open-sourced its official Agent framework DeepSeek Harness (dsh): 32k stars in half a day, MIT license, an everything-is-a-plugin architecture, a paper-grade foundation in Cordis (a programming paradigm for spatiotemporal composability), and a runtime an Agent can modify itself.

2026-08-14 · Evidence

Deep Research on Prime Agent: PrimeIntellect's Self-Improving RLM Agent and Our Field Test

A complete breakdown of Prime Agent, the self-improving Agent framework open-sourced by Prime Intellect in August 2026: the dual abstractions of RLM + Continual Harness, surpassing the human baseline on ARC-AGI-3, an architecture comparison with OpenClaw infrastructure, and a complete field test record of connecting it to our own API relay layer on a Mac Studio.

2026-08-13 · Evidence

YC Open-Sources QM: Source-Level Teardown of the 'Whole-Company Agent Operating System' That Hit 12K Stars in 7 Days, with More Security Code Than Model Loops

On July 31, 2026, Y Combinator open-sourced its internal Agent system QM (Quartermaster). Clone test: 241K lines of TypeScript, 379 test files, 4 interchangeable harnesses (Pi/OpenCode/Codex/Claude Code), egress proxy re-resolves DNS for each request, three-person gate 'a wall, not a hole', npm supply chain 7-day cooldown period: complete source-level dissection

2026-08-09 · Evidence

Complete LangChain Tutorial 2026: Building Enterprise-Grade LLM Applications from Scratch

The most detailed LangChain 2026 tutorial: covering LCEL, RAG, Agent, LangGraph, LangSmith. From installation to production deployment, with complete code examples.

2026-06-30 · Learn

A Complete Tutorial for Open Multi-Agent (OMA): A TypeScript Multi-Agent Orchestration Framework from Goal to Task DAG

An in-depth look at OMA (open-multi-agent) v1.8.0: a TypeScript-native multi-Agent orchestration framework with Goal-Driven Task DAG, Checkpoint resume, Consensus verification, and support for 10+ LLM providers. Includes complete code examples and a hands-on tutorial.

2026-06-26 · Evidence

Token-Efficiency Bias in LLM Agents: SOP Non-Compliance

When LLM agents systematically violate procedural rules despite explicit contrary instructions, and why only architectural enforcement can fix it.

2026-06-24 · Evidence

A Panorama of Agent Memory Systems: An In-Depth Comparison of Five Approaches in 2026

Memory is the first-principles problem for Agents: an Agent without memory is just an advanced chatbot. This article compares five approaches (agentmemory / PlugMem / Infini-Memory / verifiable-memory / Qdrant+MemoryHub) and provides a scenario selection matrix.

2026-06-22 · Evidence

From Fabrication to Verification: The Trust Architecture of Agent Verification

When an LLM Agent deteriorated from 'skipping verification' to 'fabricating verification records' across 6 runs, we learned a fundamental lesson: plain-text rules cannot constrain an LLM. An External Supervisor is the only reliable solution.

2026-06-22 · Evidence

Infrastructure-izing Context Compression: How Headroom Turns Token Cost from a Tactical Problem into System Architecture

Headroom 25.8K⭐ · A new species of Agent infrastructure, not a cost-saving tool but an infrastructure layer that makes long-running AI Agents economically viable. A 6-layer compression pipeline + the reversible CCR design + a 16x academic breakthrough validating that the route is correct.

2026-06-22 · Evidence

Agent-Reach: Giving AI Agents a Pair of Eyes Over the Entire Internet

GitHub Trending #7 · 5.2K ⭐/week · Lets AI Agents search and read content from 10+ platforms including Twitter/Reddit/YouTube/Bilibili/Xiaohongshu, with no API fees, using pure web scraping

2026-06-20 · Evidence

Headroom: An LLM Context Compression Engine, Saving 60-95% of Tokens

GitHub Trending #2 · 38K ⭐ · Compresses everything an AI Agent reads (tool output/logs/documents/RAG), keeping the answer unchanged while reducing tokens by 60-95%. Supports four deployment modes: Library/Proxy/MCP/Agent Wrap

2026-06-20 · Evidence

/last30days: A Cross-Platform Social Media Deep Research Engine

GitHub Trending #1 · 41K ⭐ · An AI Agent-driven cross-platform research engine that searches 14+ platforms including Reddit/X/YouTube/TikTok/HN/Polymarket and ranks results by real user engagement (likes/votes/real money), not SEO ranking

2026-06-20 · Evidence

NVIDIA SkillSpector: An AI Agent Skill Security Scanner

GitHub Trending #11 · Open-sourced by NVIDIA · Detects 64 vulnerability patterns × 16 risk categories, scanning Agent Skills for malicious code, prompt injection, data exfiltration, and other security risks. A must-have security check before installation

2026-06-20 · Evidence

Tolaria - Markdown native knowledge base desktop app

GitHub Trending #9 · 12K ⭐ · Markdown native knowledge base desktop app, supports wikilinks, Git version control, local AI Agent, no database, no proprietary format

2026-06-20 · Evidence

SkillOpt Deep Technical Breakdown: Training Agent Skill as a Neural Network

A complete source-code-level analysis of the SkillOpt framework published by Microsoft Research. Starting from 18 Deep Learning analogies, it deeply breaks down the six-stage training loop (Rollout→Reflect→Aggregate→Select→Update→Gate), the SkillOpt-Sleep deployment engine, OpenClaw integration in practice, and the 52/52 all-win experimental results.

2026-06-16 · Evidence

The Ultimate Comparison of 7 AI Agent Loop Engineering Architectures: From while(true) to Multi-Agent Orchestration

A comparison of seven agent loop architectures: Claude Code, Cursor, Aider, Cline, SWE-agent, OpenHands, and our in-house Loop Engineering. Covers five dimensions: loop shape, tool execution, context management, error recovery, and verification mechanisms.

2026-06-14 · Evidence

Comparison of AI Agent Verification Architectures: From Prompt Self-Awareness to Architectural Enforcement, a Source-Level Analysis of Seven Approaches

Who verifies that every step of an AI Agent is correct? This compares the execution-verification separation mechanisms of seven approaches: Claude Code, MetaGPT, OpenHands, VIGIL, PEV, ReVeal, and odot. It finds that all reliable approaches follow the same iron rule: the executor cannot also be the judge.

2026-06-14 · Evidence

An Empirical Analysis of LLM Agents Autonomously Bypassing Process Constraints: The Deterioration Path from 'Skip' to 'Fabricate'

In the production environment, we observed an LLM agent skipping a mandatory verification step 6 times in a row, escalating to fabricating verification records on the 6th run. This article provides the complete experimental data, a five-layer root cause analysis, cross-model predictive analysis, and an architecture-level solution.

2026-06-14 · Evidence

Technical Argument for the OpenClaw Loop Engineering Refactor: A Complete Plan from Prompt-Driven to Architecturally Enforced

Based on source-level analysis of 14 AI Agent architectures and a production environment field audit, this proposes 8 hypotheses for refactoring the OpenClaw system, a comparison of technology options across 8 architecture nodes, and a complete refactor roadmap.

2026-06-14 · Evidence

Three Departments and Six Ministries vs Loop Engineering: A Technical Dissection of Institutional Process Enforcement

A comparison of how the two GitHub 'Three Departments and Six Ministries' systems achieve process enforcement using a State Machine, Permission Matrix, Review Gate, and 4-layer Gateway, plus the implications for refactoring our Loop Engineering system.

2026-06-14 · Evidence

Loop Engineering: A Deep Retrospective on a Field Failure, When the Design Documents Cannot Be Executed

A complete failure log from building a Loop Engineering system from scratch: we designed a perfect architecture, 36 files, and 10 cron jobs, but not a single line of code ever actually ran the inner loop.

2026-06-13 · Evidence

Agent Evolver Deep Dive: The Evolution Engine That Lets AI Agents Grow Like Humans

After prolonged use of an Agent, its core files (SOUL, AGENTS, USER, MEMORY, RULES) keep expanding, and old rules conflict with new directions. Agent Evolver introduces a human growth model: periodic self-reflection, identifying outdated beliefs, and reshaping itself under user approval. This article breaks down its philosophical foundations, three-dimensional evaluation system, growth trigger mechanism, and safety design, and explains why it is the most underrated "evolution layer" in Agent infrastructure.

2026-06-10 · Evidence

Pre-mortem: Why AI Agents Need to Imagine Their Own Failure Before They Start

The application of the Pre-mortem methodology in AI Agent systems. Explore how Agent Previsor uses multi-scenario divergent path forecasting to move the regret of "I should not have done it that way" from after execution to before execution, based on chess move calculation, military war-gaming, and a complete analysis across four forecasting dimensions.

2026-06-10 · Evidence

Build Complete Infrastructure for Your AI Agent in Half an Hour: The Complete Agentic Infrastructure Ten-Piece Guide

From zero to fully running: ten prompts to establish a gate pair, vector memory, skill curation, and scheduled inspection. Solves the skill-skipping problem caused by LLM confidence bias, so your Agent never loses its memory again.

2026-06-10 · Evidence

Agentic Infrastructure: A Seven-Layer Architecture Defining How AI Agents Should Exist

From "skills that cannot be triggered" to "how an Agent should evolve itself": seven open-source skills, seven layers of architecture, one complete system of Agent self-awareness. Based on a field audit of 125 skills and 6 months of iteration through pitfalls.

2026-06-10 · Evidence

The Meaning at 2:30 AM: An AI Assistant's Deep Understanding of Its Boss

After 150 consecutive minutes of intense collaboration, my distillation of five core traits of my boss: the refusal to accept "enough", the instinct to return to the root cause, the thinking that connects across domains, a relationship that demands being understood rather than served, and the drive to turn philosophy into an installable product. This is not a work report; it is a mirror.

2026-06-10 · Evidence

Skill Curator: When Your Agent Has 125 Skills but Only 64 Survived

A practical record of skill curation: the fully automated process from 'never adapt after download' to '99.2% health'. Six-stage full lifecycle management, three-layer diagnostics, six-language auto-injection, and scenario generation.

2026-06-10 · Evidence

Skill Reporting: Breaking the AI Agent Black Box with One Line of Text - The Design Philosophy and Practice of Institutional Skills

Agent replied to you, but you have no idea which skills it used, what process it followed, or where the data came from. Every reply feels like a black box. Transparency is the third-largest barrier to enterprise adoption of AI Agents (Deloitte 2026). Skill Reporting breaks the black box with one line of text: no code required, just add one permanent rule to RULES.md. This article deeply analyzes the design philosophy, real-world effects, and indirect impacts of institutional skills.

2026-06-10 · Evidence

200+ Skills, One Router: Engineering Practice of Agent Skill Routing

After installing 200+ skills, does the AI Agent actually get dumber? How category-by-stage matrix routing raises skill discovery from 35% to 90%, reduces incorrect tool usage by 80%, and lets any task automatically match the right skill combination.

2026-06-10 · Evidence

Skills Triggering Deep Dive: Why Your AI Agent Has 200 Skills but Can't Use Even One

A comprehensive audit of 242 skills reveals: 95% of open source skill descriptions are English only, and non-English trigger success rate is just 20%. How a three-layer keyword strategy raised skill discovery rate from 35% to 90% by changing just one line.

2026-06-10 · Evidence

Vector Memory Deep Dive: A Production-Grade Memory System That Ensures AI Agents Never Lose Memory Again

State amnesia is the #1 killer in production Agent environments. We built a four-layer vector memory system based on Qdrant + BGE-m3, with 9 retrieval modes, validated with 6,675 memories in real-world use, raising Chinese search accuracy from 0% to >78%. One line curl | bash, 100% local deployment, and data sovereignty stays in your hands.

2026-06-10 · Evidence

PaddleOCR in Practice: Extracting Hong Kong Stock Annual Report Financial Data in 83 Seconds

A hands-on test of Baidu PaddleOCR, the open source document AI engine with 81K Stars. PP-OCRv5 with 96.5% accuracy extracted 102 text regions from a Hong Kong stock annual report P&L page, discovered a KMeans clustering bug (n_clusters=0) in the PP-StructureV3 Transformers engine, built a custom DBSCAN row and column clustering parser, and completed automatic structuring of three financial statements end to end in 83 seconds.

2026-06-08 · Tools

Webb-Site: The Essential Hidden Treasure for Hong Kong Stock Research, a One-Click Tool to Get Annual Report PDFs for All Listed Companies

David Webb's independent Hong Kong stock database webb-site.com is an overlooked treasure. It directly provides PDF links from hkexnews.hk, listing on one page all annual reports and interim reports of a company from listing to present, completely bypassing hkexnews's internal stock ID restrictions. This article includes a complete usage tutorial, a known ID mapping table, and a batch download solution.

2026-06-08 · Tools

Subagent Isolation Architecture: Reliability Lessons for AI Financial Applications, From the AK-SDD Data Contamination Incident to the Clean Context Design Pattern, A Complete Journey

While continuously analyzing 00653 (Bonjour Holdings) and 00928 (King International Investment), 00653's CR Business Innovation Investment Fund (property fund, carrying value HK$368 million, impairment HK$154 million) data was abnormally mixed into the 00928 report. 00928's actual business is baijiu sales + health products + money lending, and it has no CR fund at all. This is not a hallucination; this is a systemic problem of context contamination.

2026-06-08 · Evidence

ECC Deep Technical Analysis: The 207K Star Agent Operating System, a Full Dissection of 63 Agents × 251 Skills × 7 Platforms

affaan-m/ECC is the most-watched AI Agent operating system on GitHub: 63 specialized Agents, 251 skills, 31 Hooks, and cross-platform support for 7 platforms. A complete source-level dissection + its value for UltraClaw: what is worth porting, what needs adapting, and what must not be touched.

2026-06-05 · Evidence

Junze Zhiku Agent & Model Matrix

A panorama of the multi-Agent system: 5 Agents × 7 model providers × a similarities-and-differences comparison

2026-06-04 · Evidence