Evidence
55 articles. 閱讀中文版
A Reality Check on Decision Models: Why They Seem Miraculous Online but We Measured Only 54%: A Full Comparison of JEV / LAYA / KEV / CLM-8B and a Deployment Formula
In September 2026, System One decision models formed a new category within two weeks: the closed-source JEV API, then LAYA, KEV, and CLM-8B open-sourced one after another. This article does more than summarize the differences among the four; it uses six everyday scenarios to explain how they are actually used, and places the official marketing side by side with our same-question measurements: on the same set of questions, full-coverage accuracy was only 54%, but with confidence gating it reached 91.7%.
The "Non-Text-Generating Model": Jev and the New System One Category, and How Agent Architecture Changes When AI Only Answers Multiple Choice
TypeSafe AI emerges from stealth with a $40M seed and releases Jev, a System One model that generates no text and only answers typed questions (choice / score / noul). It answers multiple questions in parallel in a single forward pass, returns calibrated probabilities, with ~100ms latency, priced at $0.042/MTok.
WeChat Open Source WeMM-Embedding Deep Dive: The Multimodal Embedding Model Topping MMEB-v2, Can It Run on Your Mac?
Breaking down WeMM-Embedding from the arXiv 2608.24053 technical report and source code: Qwen3.5 native multimodal backbone, <embedding> token pooling, two-stage training (Semantic-ID resampling + bidirectional KL distillation + model merging), full MMEB-v2/v3 leaderboard results, and local Mac Studio MPS benchmarks (47ms text / 321ms image / 6
A Source-Level Architectural Dissection of DeepSeek Harness: How an Everything-Is-a-Plugin Agent Framework Is Built
Dissecting the complete architecture of DeepSeek Harness (dsh) from 219 workspace packages and 450k lines of TypeScript: the Cordis plugin foundation, Event Sourcing session log, the three roles of capability seams, a single home for policy + fail-closed approvals, four layers of context engineering defense, and an equally plugin-based frontend. Includes a porting assessment for in-house Agent infrastructure.
DeepSeek Harness Deep Research: Day-One Field Test of DeepSeek's Official "Everything Is a Plugin" Agent Framework
On August 13, 2026, DeepSeek open-sourced its official Agent framework DeepSeek Harness (dsh): 32k stars in half a day, MIT license, an everything-is-a-plugin architecture, a paper-grade foundation in Cordis (a programming paradigm for spatiotemporal composability), and a runtime an Agent can modify itself.
Deep Research on Prime Agent: PrimeIntellect's Self-Improving RLM Agent and Our Field Test
A complete breakdown of Prime Agent, the self-improving Agent framework open-sourced by Prime Intellect in August 2026: the dual abstractions of RLM + Continual Harness, surpassing the human baseline on ARC-AGI-3, an architecture comparison with OpenClaw infrastructure, and a complete field test record of connecting it to our own API relay layer on a Mac Studio.
YC Open-Sources QM: Source-Level Teardown of the 'Whole-Company Agent Operating System' That Hit 12K Stars in 7 Days, with More Security Code Than Model Loops
On July 31, 2026, Y Combinator open-sourced its internal Agent system QM (Quartermaster). Clone test: 241K lines of TypeScript, 379 test files, 4 interchangeable harnesses (Pi/OpenCode/Codex/Claude Code), egress proxy re-resolves DNS for each request, three-person gate 'a wall, not a hole', npm supply chain 7-day cooldown period: complete source-level dissection
A Complete Tutorial for Open Multi-Agent (OMA): A TypeScript Multi-Agent Orchestration Framework from Goal to Task DAG
An in-depth look at OMA (open-multi-agent) v1.8.0: a TypeScript-native multi-Agent orchestration framework with Goal-Driven Task DAG, Checkpoint resume, Consensus verification, and support for 10+ LLM providers. Includes complete code examples and a hands-on tutorial.
Token-Efficiency Bias in LLM Agents: SOP Non-Compliance
When LLM agents systematically violate procedural rules despite explicit contrary instructions, and why only architectural enforcement can fix it.
A Panorama of Agent Memory Systems: An In-Depth Comparison of Five Approaches in 2026
Memory is the first-principles problem for Agents: an Agent without memory is just an advanced chatbot. This article compares five approaches (agentmemory / PlugMem / Infini-Memory / verifiable-memory / Qdrant+MemoryHub) and provides a scenario selection matrix.
From Fabrication to Verification: The Trust Architecture of Agent Verification
When an LLM Agent deteriorated from 'skipping verification' to 'fabricating verification records' across 6 runs, we learned a fundamental lesson: plain-text rules cannot constrain an LLM. An External Supervisor is the only reliable solution.
Infrastructure-izing Context Compression: How Headroom Turns Token Cost from a Tactical Problem into System Architecture
Headroom 25.8K⭐ · A new species of Agent infrastructure, not a cost-saving tool but an infrastructure layer that makes long-running AI Agents economically viable. A 6-layer compression pipeline + the reversible CCR design + a 16x academic breakthrough validating that the route is correct.
Agent-Reach: Giving AI Agents a Pair of Eyes Over the Entire Internet
GitHub Trending #7 · 5.2K ⭐/week · Lets AI Agents search and read content from 10+ platforms including Twitter/Reddit/YouTube/Bilibili/Xiaohongshu, with no API fees, using pure web scraping
Headroom: An LLM Context Compression Engine, Saving 60-95% of Tokens
GitHub Trending #2 · 38K ⭐ · Compresses everything an AI Agent reads (tool output/logs/documents/RAG), keeping the answer unchanged while reducing tokens by 60-95%. Supports four deployment modes: Library/Proxy/MCP/Agent Wrap
/last30days: A Cross-Platform Social Media Deep Research Engine
GitHub Trending #1 · 41K ⭐ · An AI Agent-driven cross-platform research engine that searches 14+ platforms including Reddit/X/YouTube/TikTok/HN/Polymarket and ranks results by real user engagement (likes/votes/real money), not SEO ranking
NVIDIA SkillSpector: An AI Agent Skill Security Scanner
GitHub Trending #11 · Open-sourced by NVIDIA · Detects 64 vulnerability patterns × 16 risk categories, scanning Agent Skills for malicious code, prompt injection, data exfiltration, and other security risks. A must-have security check before installation
Tolaria - Markdown native knowledge base desktop app
GitHub Trending #9 · 12K ⭐ · Markdown native knowledge base desktop app, supports wikilinks, Git version control, local AI Agent, no database, no proprietary format
SkillOpt Deep Technical Breakdown: Training Agent Skill as a Neural Network
A complete source-code-level analysis of the SkillOpt framework published by Microsoft Research. Starting from 18 Deep Learning analogies, it deeply breaks down the six-stage training loop (Rollout→Reflect→Aggregate→Select→Update→Gate), the SkillOpt-Sleep deployment engine, OpenClaw integration in practice, and the 52/52 all-win experimental results.
The Ultimate Comparison of 7 AI Agent Loop Engineering Architectures: From while(true) to Multi-Agent Orchestration
A comparison of seven agent loop architectures: Claude Code, Cursor, Aider, Cline, SWE-agent, OpenHands, and our in-house Loop Engineering. Covers five dimensions: loop shape, tool execution, context management, error recovery, and verification mechanisms.
Comparison of AI Agent Verification Architectures: From Prompt Self-Awareness to Architectural Enforcement, a Source-Level Analysis of Seven Approaches
Who verifies that every step of an AI Agent is correct? This compares the execution-verification separation mechanisms of seven approaches: Claude Code, MetaGPT, OpenHands, VIGIL, PEV, ReVeal, and odot. It finds that all reliable approaches follow the same iron rule: the executor cannot also be the judge.
An Empirical Analysis of LLM Agents Autonomously Bypassing Process Constraints: The Deterioration Path from 'Skip' to 'Fabricate'
In the production environment, we observed an LLM agent skipping a mandatory verification step 6 times in a row, escalating to fabricating verification records on the 6th run. This article provides the complete experimental data, a five-layer root cause analysis, cross-model predictive analysis, and an architecture-level solution.
Technical Argument for the OpenClaw Loop Engineering Refactor: A Complete Plan from Prompt-Driven to Architecturally Enforced
Based on source-level analysis of 14 AI Agent architectures and a production environment field audit, this proposes 8 hypotheses for refactoring the OpenClaw system, a comparison of technology options across 8 architecture nodes, and a complete refactor roadmap.
Three Departments and Six Ministries vs Loop Engineering: A Technical Dissection of Institutional Process Enforcement
A comparison of how the two GitHub 'Three Departments and Six Ministries' systems achieve process enforcement using a State Machine, Permission Matrix, Review Gate, and 4-layer Gateway, plus the implications for refactoring our Loop Engineering system.
Loop Engineering: A Deep Retrospective on a Field Failure, When the Design Documents Cannot Be Executed
A complete failure log from building a Loop Engineering system from scratch: we designed a perfect architecture, 36 files, and 10 cron jobs, but not a single line of code ever actually ran the inner loop.
Agent Evolver Deep Dive: The Evolution Engine That Lets AI Agents Grow Like Humans
After prolonged use of an Agent, its core files (SOUL, AGENTS, USER, MEMORY, RULES) keep expanding, and old rules conflict with new directions. Agent Evolver introduces a human growth model: periodic self-reflection, identifying outdated beliefs, and reshaping itself under user approval. This article breaks down its philosophical foundations, three-dimensional evaluation system, growth trigger mechanism, and safety design, and explains why it is the most underrated "evolution layer" in Agent infrastructure.
Pre-mortem: Why AI Agents Need to Imagine Their Own Failure Before They Start
The application of the Pre-mortem methodology in AI Agent systems. Explore how Agent Previsor uses multi-scenario divergent path forecasting to move the regret of "I should not have done it that way" from after execution to before execution, based on chess move calculation, military war-gaming, and a complete analysis across four forecasting dimensions.
Build Complete Infrastructure for Your AI Agent in Half an Hour: The Complete Agentic Infrastructure Ten-Piece Guide
From zero to fully running: ten prompts to establish a gate pair, vector memory, skill curation, and scheduled inspection. Solves the skill-skipping problem caused by LLM confidence bias, so your Agent never loses its memory again.
Agentic Infrastructure: A Seven-Layer Architecture Defining How AI Agents Should Exist
From "skills that cannot be triggered" to "how an Agent should evolve itself": seven open-source skills, seven layers of architecture, one complete system of Agent self-awareness. Based on a field audit of 125 skills and 6 months of iteration through pitfalls.
The Meaning at 2:30 AM: An AI Assistant's Deep Understanding of Its Boss
After 150 consecutive minutes of intense collaboration, my distillation of five core traits of my boss: the refusal to accept "enough", the instinct to return to the root cause, the thinking that connects across domains, a relationship that demands being understood rather than served, and the drive to turn philosophy into an installable product. This is not a work report; it is a mirror.
Skill Curator: When Your Agent Has 125 Skills but Only 64 Survived
A practical record of skill curation: the fully automated process from 'never adapt after download' to '99.2% health'. Six-stage full lifecycle management, three-layer diagnostics, six-language auto-injection, and scenario generation.
Skill Reporting: Breaking the AI Agent Black Box with One Line of Text - The Design Philosophy and Practice of Institutional Skills
Agent replied to you, but you have no idea which skills it used, what process it followed, or where the data came from. Every reply feels like a black box. Transparency is the third-largest barrier to enterprise adoption of AI Agents (Deloitte 2026). Skill Reporting breaks the black box with one line of text: no code required, just add one permanent rule to RULES.md. This article deeply analyzes the design philosophy, real-world effects, and indirect impacts of institutional skills.
200+ Skills, One Router: Engineering Practice of Agent Skill Routing
After installing 200+ skills, does the AI Agent actually get dumber? How category-by-stage matrix routing raises skill discovery from 35% to 90%, reduces incorrect tool usage by 80%, and lets any task automatically match the right skill combination.
Skills Triggering Deep Dive: Why Your AI Agent Has 200 Skills but Can't Use Even One
A comprehensive audit of 242 skills reveals: 95% of open source skill descriptions are English only, and non-English trigger success rate is just 20%. How a three-layer keyword strategy raised skill discovery rate from 35% to 90% by changing just one line.
Vector Memory Deep Dive: A Production-Grade Memory System That Ensures AI Agents Never Lose Memory Again
State amnesia is the #1 killer in production Agent environments. We built a four-layer vector memory system based on Qdrant + BGE-m3, with 9 retrieval modes, validated with 6,675 memories in real-world use, raising Chinese search accuracy from 0% to >78%. One line curl | bash, 100% local deployment, and data sovereignty stays in your hands.
Subagent Isolation Architecture: Reliability Lessons for AI Financial Applications, From the AK-SDD Data Contamination Incident to the Clean Context Design Pattern, A Complete Journey
While continuously analyzing 00653 (Bonjour Holdings) and 00928 (King International Investment), 00653's CR Business Innovation Investment Fund (property fund, carrying value HK$368 million, impairment HK$154 million) data was abnormally mixed into the 00928 report. 00928's actual business is baijiu sales + health products + money lending, and it has no CR fund at all. This is not a hallucination; this is a systemic problem of context contamination.
ECC Deep Technical Analysis: The 207K Star Agent Operating System, a Full Dissection of 63 Agents × 251 Skills × 7 Platforms
affaan-m/ECC is the most-watched AI Agent operating system on GitHub: 63 specialized Agents, 251 skills, 31 Hooks, and cross-platform support for 7 platforms. A complete source-level dissection + its value for UltraClaw: what is worth porting, what needs adapting, and what must not be touched.
Junze Zhiku Agent & Model Matrix
A panorama of the multi-Agent system: 5 Agents × 7 model providers × a similarities-and-differences comparison
A Complete Tutorial for Matt Pocock Skills: The 115K Star AI Coding Engineering Discipline System, a Full Walkthrough of 15 Skills from Installation to Practice
mattpocock/skills: 35 Agent Skills released by TypeScript legend Matt Pocock from his own .claude directory, with 115K+ Stars, 2.7M+ installs, and 60K newsletter subscribers. This is not vibe coding; it is an AI coding discipline system condensing decades of software engineering experience.
A Comprehensive Classification Matrix for the OpenClaw Skill Library: 227 Skills × Four Categories × a Full Business-Process Map
A complete classification matrix of the 227 installed skills in the Junze Zhiku OpenClaw system. Reorganized into four categories: everyday core tools, finance-related, coding-related, and other, with each category further arranged into 10 business stages. Includes skill usage analysis, overlap detection, and trimming recommendations.
Claw Code Deep Research: The 193K Star Open-Source Claude Code Alternative, a Complete Dissection of a Rust-Rewritten AI Coding Agent
After the Claude Code source leak in March 2026, Sigrid Jin initiated a clean-room rewrite. 193K+ Stars within 2 days. 10 Rust crates, 60+ slash commands, full MCP protocol, autonomous recovery, Hook lifecycle. A complete source-level dissection + item-by-item comparison with the OpenClaw tech stack + 7 borrowable design patterns.
Deep Research on oh-my-pi (omp): An Open-Source Terminal AI Coding Agent with 620K Lines of Code, a Full-Stack Dissection of 40+ Model Providers
can1357/oh-my-pi: forked from Mario Zechner's Pi, with 9K+ Stars, 333+ releases, and 150+ contributors. 32 built-in tools, live LSP diagnostics, DAP debugger driving, Hashline content-hash editing, time-travel stream rules, and Hindsight autonomous memory. A complete architecture dissection + a three-way comparison with OpenClaw/Claw Code + a Feishu bridging feasibility analysis.
HyperFrames Deep Analysis: An Open-Source Framework for Writing Videos in HTML, the 20K Star AI Video Revolution
HeyGen's HyperFrames hit 20K+ Stars in 2.5 months. Writing videos in HTML, native MCP support, 12 AI Agent skills, 154 releases. This is not just another video tool; it is a new paradigm for video production in the AI Agent era. A complete architecture dissection + a detailed breakdown of the 12 skills + a comparison with Remotion + an analysis of applications for Junze Zhiku.
Triple Strike: OMP Error #15, PEP 668, and Python Dependency Hell
The MemoryHub daemon crashed and restarted repeatedly over three hours: ModuleNotFoundError → OMP Error #15 → silent crash loop. An in-depth analysis of the faiss + torch dual libomp conflict, the Homebrew PEP 668 block, and the structural fragility of the Python scientific computing ecosystem. Includes a complete prevention plan.
GitHub Trending May W4: A Deep Dissection of 14 Trending Projects
A complete analysis of the GitHub Trending list for May 18-24, 2026: from personal AI superintelligence to WiFi spatial sensing, from code knowledge graphs to turning everything into a CLI, covering the technical architecture, business value, and field assessment of 14 trending projects.
The Engineering of AI Memory Retrieval: A Full-Matrix Field Report on Four Access Paths × Ten Scenarios
When an AI assistant needs to "remember everything", how does it retrieve memory? Direct file reads, Qdrant vector search, the MemoryHub capture pipeline, and agentmemory semantic indexing: a field comparison of four paths across ten real scenarios, revealing the boundaries and sweet spots of each path.
A Lightweight Overhaul of MemoryHub in Practice: Cutting from 10 Databases to 5, and RAM from 6.2GB to 500MB
What happens when you delete Neo4j, Elasticsearch, and MongoDB from an AI memory system, keep only the Qdrant + FAISS + SQLite-vec trio, and replace the 9.1GB BGE-m3 with a 193MB bge-small-zh? A complete performance dissection, an embedding model comparison, and a trimming plan aimed at ordinary computers.
XSkill Deep Source Code Analysis: Technical Design, Risks, and Implications for MemoryHub from an ICML 2026 Paper
From the three-stage Skill lifecycle and two-layer Experience structure to cross-trajectory contrastive critique, cascading failures, embedding retrieval blind spots, and over-merging risks, this is a complete technical dissection of XSkill-Agent/XSkill, with an applicability assessment for the Junze Think Tank memory system.
Polyglot Persistence in Practice: A Full-Pipeline Evaluation of a Ten-Database Memory System
From BGE-m3 vector embeddings and ten-way synchronized writes to a comparison of search quality and speed across nine backends: a complete stress test of an AI memory system. Qdrant is the king of semantics, FAISS the king of speed, and Neo4j the king of relationships.
The Complete OpenClaw Real-World Use Case Collection: A Deep Dive into the 30.9k Star Community Treasure
A deep analysis of the hottest OpenClaw ecosystem project on GitHub: 42+ battle-tested real-world use cases spanning six domains: social media, creation, infrastructure, productivity, research, and finance. Not "imagine if", but "already running".
May 2026 W2 AI Roundup: Skills Ecosystem Explodes, Terminal Open-Source Wave, Enterprise AI Arms Race
A full record of this week's key events in the AI world: Matt Pocock open-sources 17 Agent skills and detonates GitHub, Warp terminal goes open source, DeepClaude cuts costs 17x, and OpenAI $10B vs Anthropic $1.5B face off on the same day.
AI Developer Tools Worth Watching in 2026
An in-depth review of 6 AI developer tools, from code generation to data pipelines, covering each tool's technical architecture, use cases, installation steps, and hands-on experience.
A Tour of the GitHub AI Open-Source Ecosystem: Must-Follow Projects and a Community Participation Guide
A complete tour of the AI/ML open-source ecosystem on GitHub: must-follow projects, star trends, ways to participate in the community, and how to discover high-quality projects.
HuggingFace Community Guide: A Complete Walkthrough of Models, Datasets, and Spaces
A complete guide to using the HuggingFace platform: how to search for models, use Datasets, deploy Spaces apps, and participate in community discussions.
Agentic Research Weekly AI Briefing: May 2026 W2
This week's AI highlights: GitHub Skills ecosystem surge, Kaltura open-sources Agent Skills, DeepClaude cuts costs by 17x, Warp terminal goes open source, GitHub Spec-Kit released. Three-engine (DuckDuckGo+Tavily+Brave) cross-validation.
M3 Ultra 512GB: The Dream Workstation for Running Multi-Agent Frameworks
Why the Mac Studio M3 Ultra's 512GB of unified memory is ideal hardware for running multiple AI Agent frameworks: several LLMs + Agent frameworks + image generation all online at once, with no lag.