Agentic Research

A Source-Level Architectural Dissection of DeepSeek Harness: How an Everything-Is-a-Plugin Agent Framework Is Built

2026/08/1680 min readBryan Chan閱讀中文原文
TopicsDeepSeek HarnessCordisAgent FrameworkArchitecture DesignOpen Source

The One-Sentence Version

DeepSeek Harness (dsh) is not an agent; it is the "skeleton" that runs agents. Model adaptation, the tool system, session management, sandboxing, approvals, the Web UI, and the plugin mechanism are all built in, and every part can be replaced through configuration. Its architectural philosophy is one sentence: Everything is a plugin; there is no privileged core that needs patching, and even the agent loop itself is a plugin.

This is a source-level dissection of github.com/deepseek-ai/deepseek-harness (v0.1.0-rc.5, developer preview). It is not a feature overview; it looks at how it builds replaceability into its bones.


Project Scale

ItemValue
Organizationpnpm monorepo
Workspace packages219 @deepseek-ai/dsh-*
TypeScript files~2,046 files / ~450k lines
Native code~300 lines of C11 (Linux sandbox launcher landlock-run)
Subsystem docs40+ bilingual (en/zh paired validation)
Foundation frameworkvendored Cordis 4.0.0-rc.7
LicenseMIT

It runs with a single command:

npx @deepseek-ai/dsh web   # Start Web GUI, default http://127.0.0.1:3080

1. The Foundation: The Cordis Plugin Framework

DSH does not bring in an external plugin framework. Instead it vendors Cordis in its entirety (renamed into the @deepseek-ai scope), fully owning its own framework layer, which can be audited, patched, and version-locked. Cordis's five core concepts are the key to understanding everything:

ConceptMeaning
A plugin is a ServiceA plugin is a function with inject + apply(ctx), or a Service subclass, whose lifecycle is attached to the context
The context is a service containerServices occupy stable ctx.<key> slots (ctx.tools, ctx.llm, ctx.sessions), looked up by key rather than by importing an implementation
inject declares dependenciesA plugin waits for its dependency services to be ready before starting; load order is driven by the dependency graph, with no manual orchestration
Typed eventsEvent names are registered via TS declaration merging; four dispatch modes: emit (observe), waterfall (wrap middleware, next() can delegate or short-circuit), parallel (concurrent), serial (ordered)
Reversible side effectsEverything is installed via ctx.effect() / ctx.on(), and is automatically undone when the plugin unloads

Because even the agent loop itself is a plugin, along with model adapters, the tool registry, session logs, and CLI flag parsing, every part of the product can be replaced. This is the bedrock of the entire design.


2. Monorepo Layering

vendor/      vendored Cordis framework layer (cordis/loader/hmr/include/timer/schemastery/cosmokit)
packages/    219 @deepseek-ai/dsh-* packages, divided into 40+ groups by <group>/<pkg>
  core/        product API backbone: session, system-prompt, tools, agent, agent-loop, scope
  llm/         LLM vocabulary + adapters (deepseek, pi-ai) + retry + token-meter
  api/ typert/ remote API gateway + type graph generator
  fs/ shell/ subprocess/ sandbox/ terminal/ lsp/ jobs/ web/ skill/ ...   capability seam family
  subagent/ workflow/ goal/ plan/ compaction/ spill/                    orchestration and context engineering
  session/ storage/ session-query/                                      persistence data plane
  interaction/ guard/ preset/ hooks/                                    human-machine collaboration and guardrails
  bundle/      installable profile patch layer (base / web-app / headless)
  host/ client/ Web GUI backend half and browser half
  boot/        app-bin startup glue
apps/        cli (dsh command) + web (Vite frontend entry)
native/      landlock-run (Linux sandbox launcher)
python/      Python SDK and single exe distribution runtime
examples/    runnable cordis.yml example leaves

Dependency discipline is key: extension plugins depend only on a Service Definition, never on a concrete Provider. dsh-agent-loop can be swapped, while the UI and tools depend only on dsh-agent. The dependency graph is generated from peerDependencies (docs/module-graph.md), and CI guards its freshness.


3. Assembly Mechanism: Profile / Bundle / Patch

A running dsh is a plugin tree composed of layers stacked in order at startup:

  1. Bundle: a Cordis configuration + mount code distributed via npm. dsh-base is the first layer of every profile (model adapters, tools, persistence, sandbox and approval policy, settings, credentials, telemetry); dsh-web-app adds the browser application; dsh-headless adds a one-shot runner.
  2. Profile: a named assembly in the Harness home, listing the stacked bundles plus the user's own cordis.patch.yml.
  3. Patch: locates an entry by id and replaces its config wholesale, or inserts a new entry. Stacking order: bundle (in profile order) → profile patch → home-level patch → --patch overlay, with later writes overriding earlier ones.

A few elegantly designed details worth savoring:

  • Flags as services: --host/--port are parsed by an ordinary web-startup plugin into a webStartup service, and the webserver entry lazily interpolates with !!js ctx.webStartup.port ?? 3080. commander only recognizes the two launcher flags --profile/--patch; everything else is passed as-is to the plugin tree.
  • Symlink field: composeProfile maintains a symlink field at $DSH_HOME/profiles/node_modules, guaranteeing that the dependency closure shares a single cordis instance, so the plugin graph never ends up with duplicate framework instances.
  • Hot reload: after startup, watchUserPatches hot-reloads the user patch layer via HMR.
  • Dumpable: dsh --profile web --dump-config prints the actual config tree, and any entry in it can be replaced by a user patch.

The roughly 60 insert lines of dsh-base cover: models (llm, llm-deepseek, llm-pi-ai, llm-retry), persistence (session, session-persistence-jsonl, session-projection, session-query-sqlite), sandbox/approval (sandbox-local, sandbox-policy defaulting to workspace-write, platform-exclusive bash/pwsh-sandbox, approval, permission-presets), settings/credentials/telemetry (otel defaulting to DISABLED), and all model-facing tools (bash/fs/web/skill/todo/goal/subagent/workflow/ralph/jobs).


4. Core Runtime

Agent and Agent Loop: Interface and Implementation Separated

  • The Agent handle (core/agent): the public interface; send/followup/steer/inject unify delivery; the ctx.agents registry uses setFactory() to delegate creation to the agent-loop, separating interface from implementation so the loop can be swapped.
  • ReactLoopAgent (core/agent-loop): a single-phase state machine of idle | maintenance | running, with wakeDriver → kick → while(turn()).
  • Inbox: dual queues (next-turn / next-step), and it is a persistent projection: every change first lands in the log as an agent/inbox/spliced event before the in-memory change, and on resume it is replayed from the seed boundary.

Turn / Step Flow

A step = one model request plus the tools it calls; a turn = 0..n steps:

turn/start
  claim next batch of input
  → agent/pre-step (waterfall: can reject/rewrite)
    step/start → user/message sunset log
    system-prompt/assemble (waterfall)
    → agent/request → llm/stream → assistant/chunk* → assistant/message
    → tool/call* → tools/pre-execute → tools/execute → tools/post-execute → tool/result*
    step/end → (work still owed? claim next batch → next step)
  → agent/turn-stopping (serial)
turn/end

Tool dispatch has an exclusive barrier + a bounded parallel rolling pool, and results are submitted in model order; on request failure it goes to agent/request-error (which can decide to retry).

Session Log: Event Sourcing for Everything

The append-only SessionEvent log is the single source of truth. deriveMessages() projects the model history from the log; raw assistant/chunk events guarantee replay and UI fidelity. Fork, resume, transcript, telemetry, and persistence are all derived from this single event stream.

The core invariant: "what the model can see is already recorded"; everything reaching a model request must be reconstructible from the log (asserted by a runtime invariant). request/header (an EpochHeader snapshot) makes every request a pure function of the log. On append it performs lossless-JSON validation + deep freezing, with a continuous seq.

Tool System and Execution Pipeline

  • ToolDefinition = schema + mandatory output declaration + execute + finalizeContent + present*; the defineTool DSL performs type inference.
  • The registry is layered by scope + ToolRestriction (allow/deny), and schemas() projects the allowlist into the prompt.
  • The pipeline: tools/pre-execute (allow/deny/ask) → ask is gated once via ctx.approval (unanswerable means deny) → monotonic ToolGuard (only deny/abstain, and order cannot be overruled) → tools/execute (wrapped with timeout/retry) → the tool body → tools/post-execute → normalization → tool/result logged.

Three Event Domains

Event DomainNaturePurpose
Session events (turn/*, assistant/*)Persistent facts, appended to diskFacts that must survive a reload
Agent events (agent/*)Real-time, carrying the live AgentObserving and intercepting work in progress
Capability events (fs/*, tools/*, llm/stream)Hook points on seamsAttaching policies and adapters without import cycles

5. The Capability Seam Pattern: How the Execution Layer Is Organized

Every replaceable capability consists of three roles:

  • Service Definition: declare module merges ctx.<name> into the Context + an abstract Service subclass, with zero implementation in the package itself;
  • Service Provider: implements and registers (a single-implementation seam throws if a second is loaded; registry-style seams such as subagent/llm coexist by name);
  • Consumer: only injects the ctx key, importing no provider whatsoever.

Switching providers = swapping one plugin in cordis.yml, with model-facing tools and the agent loop entirely untouched.

Its most brilliant application is the execution world: the contracts of ctx.fs and ctx.subprocess are bound to "the same world". The bash executor, the PTY backend, the LSP host, and the external CLI subagent all depend only on ctx.subprocess + ctx.fs. So after swapping those two providers for e2b implementations (fs-e2b + subprocess-e2b sharing one E2B SDK handle), bash/PTY/LSP/claude-code and the rest move as a whole into a remote Linux sandbox, with not one line of consumer code changed. The unit of migration is "the world", not a single tool, which is the upper bound of the seam pattern's payoff.

Sandbox and Approval

  • ctx.sandbox: process fencing within the same world. The consumer hands over the exact argv it is about to spawn, and the provider returns a ConfinedArgv + enforcement: full|partial per the per-call policy. Backends: Linux bwrap / landlock (a native C launcher, fail-closed) / macOS seatbelt / Windows private SID+ACL.
  • ctx.sandboxPolicy is the single home for policy: the two enforcement families, bash and fs, read the same parsed result and the same writableRoots, structurally eliminating fence drift; the parsed result is written into the session log, and replay rebuilds the same policy.
  • Escalation choreography: escalation.ts unifies the widening mode ladder, pairing-validation of sandbox_permissions ⇔ justification, and approveEscalation goes through ctx.approval fail-closed before any execution, then retries with a one-time wider policy on approval.
  • Approval seam: approval/request waterfall dispatches a one-time decision, and when there is no responder it fails closed to unavailable; approval/asked|decided are logged for audit.
  • Honest positioning: fs-sandbox and the workflow vm both explicitly document themselves as containment rather than a security boundary; true kernel-level isolation is left to the sandbox runner, and the enforcement level is reported truthfully.

6. Multi-Turn Orchestration: subagent / workflow / ralph / goal

MechanismNatureSuitable For
subagentA ctx.subagents named-provider registry: in-process spawn/fork, external CLI (claude-code/codex), the ACP protocol, dsh-sdk. Capabilities are declared with static descriptors, and unsupported ones are explicitly rejected with UNSUPPORTED_CAPABILITY, never silently degradedDelegating independent subtasks
workflowctx.workflowEngine runs a model-written JS orchestration script in a worker-thread vm context; within the script, agent() bridges to subagents for fan-out, and phase()/log() report progressLarge-scale fan-out orchestration
ralphA frozen usage on the workflow engine: each round starts a brand-new subagent, carrying only an immutable objective + a bounded handoff (16KB) between rounds, sharing the workspace as long-term memoryIterating with fresh context (isolating contamination)
goalA same-session mechanism: the goal state is folded out of the session log (with revisions), and the round driver enqueues subsequent rounds into the same Agent's inbox without spawning a subagentAutonomous continuation within a session (preserving context)

There is also ctx.jobs (a general background job runtime + the job_* tools), ctx.schedule (in-session scheduled follow-ups), and plan-mode (collaborative planning state).


7. Context Engineering: Four Layers of Defense

A complete context window defense line, made up of four seamlessly connected mechanisms:

  1. token-meter: per-session replay-fold token metering, to gauge pressure;
  2. compaction-tool-result-pruner: first rewrites an oversized current tool result into a replayable replacement node (no model needed);
  3. compaction: LLM summarization. The three events compaction/start|summary|end only write to the log (lock, summarize, release; a mid-way crash manifests as a detectable leftover lock), and the summary makes its only surface change through a user/message with surfaceOp: {op:'replace'};
  4. spill: a tools/post-execute transformer; when a result exceeds the threshold, the full text is stored in the spill store, and the model receives only a bounded head/tail preview + a locator; failure never turns a successful call into an error.

8. Persistence and Projection

  • Persistence: the SessionEvent event-source log is the only authority; JSONL (one file per session) and SQLite (node:sqlite) share a PersistenceCoordinator (batched writes, prepared-session cache).
  • Projection: domain plugins contribute pure synchronous ProjectionDefinitions (init/apply/view), and the framework maintains read models by folding committed events. The whole-value event rule (an event must carry the full post-state; no bare deltas) keeps folding cheap and makes unit transfers self-describing.
  • Cache: projection checkpoints are persisted, and cold reads walk the ladder of "cache row + persistence tail replay", so a list page never loads the full log.
  • Relationship: persistence is the authoritative log, projection is the folded read model, and cache is a shortcut rather than an authority.

9. Web GUI: Frontend and Backend Are Equally Plugin-Based

Tech stack: React 18 + TypeScript + Vite (apps/web is only an entry point; rejectStandaloneServe forbids running it bare, and only the host can inject window.__DSH_BOOT__); CSS Modules + --dsw-* design tokens, with no Tailwind or component library.

Communication protocol: simplex RPC (POST /api/<namespace>/<method>) + two downstream-only WebSockets (/api/events.mux for the session event multiplexed stream, /api/events.host). Trust boundary: /api uniformly checks Host/Origin/sec-fetch-site up front (preventing DNS rebinding); privileged methods (directory selection, settings, credentials) are restricted to loopback.

Three key designs of a plugin-based UI:

  • window.__DSH_BOOT__: the plugin entry graph that the host injects via index-tap. The shell starts in two phases: a loading page → the cordis Loader fully ACTIVE → a one-time switch to the real UI. The main bundle contains only the shell and the module system; every functional UI is a plugin bundle loaded at runtime.
  • The Slots mechanism: a single table register() for "declaration = render grant = runtime contract"; chain slots invert the choice (entries self-nominate, and the first non-null wins), so preemptive UIs such as approval/question need no central switch.
  • ConversationNodeDefinition: chat messages use an open registration system; plugins declaration-merge the keys of ChatNodeDataMap and register a match/start/update/buildViewNode state machine + a keyed renderer; events are replay-foldable by seq.

Typert: a strict build-time contract. Business services are marked with the @Remote decorator; at build time a type graph / Zod schema / Remote descriptor is generated from the host's ts.Program, and the client only mounts the artifacts. Complex host objects (such as Agent) are swapped for wire ids via lookup; cold sessions resume automatically.


10. LLM Abstraction Layer and MCP

ctx.llm is the message and streaming vocabulary (Message/ContentBlock/StreamChunk) plus an adapter registry. Providers: llm-deepseek (official), llm-pi-ai (a multi-provider directory); llm-retry provides a retry strategy in plugin form. Adding a model vendor = registering an adapter with ctx.llm, a matter of one line of configuration. MCP integration: mcp-client connects one MCP server per instance, and tools are registered into ctx.tools as mcp__<server>__<name>.


11. Engineering Practice and Quality System

  • Docs as code: 40+ subsystem documents maintained bilingually (en/zh paired validation); module-graph.md, capability-seams.md, and the configuration catalog are generated from source by scripts with CI freshness guards; type-equiv code blocks are checked for drift against source.
  • Test matrix: vitest unit tests + e2e + snapshots (recording/replaying LLM interactions) + web stress and performance tests + a Windows (Wine) gate; llm-mock-server/llm-replay let tests avoid depending on the real API.
  • Invariant registry: ctx.invariants lets packages register their own runtime invariants (such as "what the model can see is already recorded"), which the framework asserts centrally.
  • Pre-release posture: explicitly "foundation before compatibility", free to rename and repackage, using a monotonic SCHEMA_VERSION for SQLite, carrying no compatibility baggage.
  • Agent-friendly: the repo ships with AGENTS.md/CLAUDE.md, .agents/notes/ (one Agent Note per design decision), and a step-by-step cookbook guide. The project itself is designed to be maintained by agents.

12. Summary Table of Key Design Decisions

#DecisionBenefit
1Everything is a plugin, no privileged core (even the agent loop is a plugin)Any part can be replaced from configuration; registrations are undone on unload
2Event-source everything (history/inbox/request headers all derived from SessionEvent)Fork/resume/replay/audit share one source; "what the model can see is already recorded" is assertable
3Three roles of capability seams + the execution world boundarySwapping 2 providers migrates the entire execution world to a remote sandbox
4A single home where policy and execution are separated (sandboxPolicy) + fail-closed approvalStructurally eliminates fence drift; absence of approval = denial
5Three event domains + four dispatch modes (waterfall interception points + monotonic guard)Policies can be stacked, and order is safe
6Declaration merging extends the core sum typesPlugins extend core types without changing the core, all checked at compile time
7A self-sufficient shell + a runtime UI plugin graph + Typert build-time contractsFrontend and backend are equally plugin-based; cross-process RPC is fully type-safe
8Layered context engineering defense (token-meter→pruner→compaction→spill)Long sessions do not blow the window, and every layer is replayable and auditable
9The framework layer is vendored and ownedAuditable, patchable, version-locked
10Whole-value event rule + projection/cache derivationCheap folding, cold reads walk a ladder; there is only one authority
11Patches replace config wholesale + flags as services + a symlink fieldAt most one owner per line across bundle layers and the user layer; predictable, dumpable, hot-reloadable

13. Porting Assessment for In-House Agent Infrastructure

After reading this architecture, what is most worth stealing is not any particular feature, but four structural principles. Using our own multi-Agent infrastructure (the UltraClaw / OpenClaw ecosystem) as the reference:

① The "what the model can see is already recorded" invariant → root-curing inconsistent subagent data reporting

A pit we hit: an AK-SDD subagent mixed 00928's data into the 00653 report, and its manual reporting dropped items. The root cause is that there is no assertable bridge between "what the subagent says it did" and "what it actually did". dsh's solution is to force everything reaching the model to be reconstructed from an append-only log. Portable practice: force every subagent's output to write a structured event (input + tool calls + result hash), and have the main agent verify against the event stream rather than the subagent's self-report. This dovetails exactly with our existing external supervisor approach.

② A single home for policy → eliminating rule fence drift

Our current state: behavioral rules are scattered across PERMANENT-RULES / RULES / AGENTS / SOUL, the same constraint may have multiple formulations, and changing one place misses another. dsh's sandboxPolicy is the single home for policy, and both bash and fs read the same parsed result. Portable practice: converge "which directories are writable, which commands require approval, which sources are banned" into one machine-readable policy file (YAML/JSON) that every skill and cron reads, rather than each hardcoding its own.

③ Fail-closed approval → upgrading our Gate mechanism

dsh's approval seam fails closed to denial when there is no responder, and approval decisions are logged for audit. Our Gate/verify.py already has a prototype, but two things can be added: defaulting to denial when approval is absent (rather than defaulting to allow or hanging), and writing every Gate decision into a traceable event stream.

④ Four layers of context engineering defense → confirming headroom's direction, adding spill

We have already deployed headroom for context compression. dsh's additional inspiration is spill: when a tool result exceeds the threshold, give the model only a head/tail preview + locator, store the full text in a store for retrieval on demand, and "failure never turns a successful call into an error". This suits our hundred-KB annual report and PDF extraction scenarios better than plain compression, and it can become a standard post-processing step in the AK-OCR pipeline.

What we do not recommend porting directly for now: the entire Cordis plugin foundation and Typert build-time contracts. These are framework-level refactors with big payoffs but very high migration costs, better treated as long-term references than short-term ports. At our current stage, applying principles ①②③④ to the existing infrastructure offers the best value.


Conclusion

What moved me most about DSH is its honesty: fs-sandbox explicitly states that it is "containment rather than a security boundary", the pre-release explicitly states that "there will be breaking changes", and unsupported capabilities explicitly return UNSUPPORTED_CAPABILITY rather than silently degrading. A framework that dares to write its own boundaries into the documentation is more convincing than any marketing language.

And "everything is a plugin" is not a slogan: when even the agent loop, CLI flags, and frontend UI are made into unloadable plugins, it becomes an architectural discipline that can be verified from the source code.


Analysis method: clone the source code, have multiple parallel subagents read the core runtime, boot/plugin system, Web GUI, and execution layer in depth, and cross-validate against the official architecture documentation. This article is the source-level sequel to the 2026-08-14 "DeepSeek Harness Deep Research: Day-One Field Test".