Agentic Research

DeepSeek Harness Deep Research: Day-One Field Test of DeepSeek's Official "Everything Is a Plugin" Agent Framework

2026/08/1460 min readBryan Chan閱讀中文原文
TopicsDeepSeek HarnessCordisAgent FrameworkOpen SourceDeep Research

The One-Sentence Version

At 19:56 (HKT) on August 13, 2026, DeepSeek officially open-sourced its own agent framework, DeepSeek Harness (dsh), which surged to 32,000+ stars within half a day. It is MIT-licensed, a TypeScript monorepo, and its core idea is "everything is a plugin": model adapters, the tool registry, the session log, and even the agent loop itself are all plugins, every one of them replaceable from configuration, with no privileged core that needs patching.

The foundation is the paper-grade plugin framework Cordis (with the companion paper "A Programming Paradigm for Spatiotemporal Composability" released as a draft the same day). The most radical capability: an Agent can define and mount new plugins at runtime to modify itself (tool-cordis; the official demo is literally called "the agent modifies its own runtime").

I completed a source clone, installation, Web UI startup, and a headless end-to-end field test on the night of release (connected to our own API gateway, running DeepSeek V4 Pro successfully). This article is the complete record.


Release Data

ItemValue
Created2026-08-13 19:56 HKT
Stars (half a day)32,061
Forks2,407
Version0.1.0-rc.6 (npm @deepseek-ai/dsh)
LanguagesTypeScript (plus a Python SDK and a C11 native sandbox)
LicenseMIT
StatusDeveloper Preview (officially stated to include breaking changes)

One-line launch:

npx @deepseek-ai/dsh web    # Web UI at http://127.0.0.1:3080

Core Philosophy: Everything Is a Plugin

Most agent frameworks are architected as "one core loop plus a pile of extension points". dsh goes further:

On top of Cordis, every part of the product is a plugin, including model adapters, the tool registry, the session log, and the agent loop itself. So every part can be replaced from configuration. There is no privileged core to patch: you extend dsh by mounting a plugin alongside other plugins, and registration itself is a "revertible effect" that automatically rolls back when the plugin unloads.

This means "replacing the agent loop" in dsh is not forking the source code, but changing one line of configuration.

The Five Cordis Concepts Behind This Philosophy

  1. A plugin = a Service: an object implementing a Service (a function with an optional inject + apply(ctx), or a Service subclass), whose lifecycle Cordis mounts into the context.
  2. Context = a service repository: services occupy stable ctx.<key> slots (ctx.tools, ctx.llm, ctx.sessions, and so on); plugins find services by key rather than importing a concrete implementation.
  3. Dependencies are declared via inject: after a plugin declares the services it needs, it waits until they exist before activating; load order is expressed by service dependencies, with no manual startup ordering needed.
  4. Typed Events: a public contract made of four dispatch modes: emit (observe, do not wait), waterfall (a middleware chain, with a return value), parallel (concurrent), serial (sequential, awaiting).
  5. Registration = a revertible effect: prompt sections, tool schemas, adapters, and listeners are all installed via ctx.effect() / ctx.on(), so reload and teardown roll back predictably.

Cordis: A Paper as the Foundation

dsh did not just "use a plugin library"; it built its entire framework layer on top of a formal paper: "A Programming Paradigm for Spatiotemporal Composability" (draft of 2026-08-13, released the same day as dsh).

The paper splits the problem of "dynamic composition" into two orthogonal dimensions:

  • Temporal composability: when a component is removed, its side effects can be fully undone. The solution elevates the classic effect concept into a runtime mechanism, revertible effects: every context transformation carries an inverse transformation, tracked by the runtime.
  • Spatial composability: declarative and reactive management of dependencies between components. The solution is reactive coeffects: every change to the context notifies a component according to its coeffect specification.

The two are unified into a single context type, forming a programming paradigm; combining them yields the component concept and a calculus of dynamic composition, whose metatheory generalizes composability from a single component to a system composed in an interleaved way. Cordis is the implementation of this paradigm: a core library (effect tracking + coeffect resolution) plus a declarative component loader (config reconciliation + HMR).

One notable detail: dsh vendors the Cordis source code (a source-level copy) into the monorepo, renamed into the @deepseek-ai scope (cordis 4.0.0-rc.7 + cosmokit + schemastery, etc.), on the grounds that the framework layer must be "auditable, patchable, and version-pinned". Releasing dsh therefore also releases the entire framework layer it owns.


Deep Dive into the Architecture

1. Profile and Bundle: A Plugin Tree Composed at Boot

One dsh run is a plugin tree composed in layers at boot:

  • Profile: a named composition (web and headless are built-in templates), listing the bundles it stacks + out-of-tree plugins + the user's own cordis.patch.yml.
  • Bundle: the distribution format of Cordis config lines plus their mount code. dsh-base (model adapters, tools, persistence, sandbox and approval policy, settings, credentials, telemetry) is the first layer of every profile; dsh-web-app adds the browser application; dsh-headless adds a one-shot runner.
  • Stacking order: the bundles listed by the profile → the profile's patch → home-level patch → --patch overlay. Any config line can be replaced wholesale by a higher layer's patch.

dsh --profile web --dump-config prints the actual tree booted on your machine, and every line of it can be replaced by your patch.

2. Session Log: Model-visible ⟺ logged

The session log is the single source of the context the model sees. deriveMessages() projects the model history from the log, and raw assistant/chunk events preserve replay and UI fidelity. Fork, resume, transcript, telemetry, and persistence are all derived from this one stream.

The runtime invariant: anything reaching a model request must be reconstructible from the log, so any new model-visible input must be a new session event (extend SessionEventMap and render from the log).

3. Turn / Step Flow

Step = one model request plus the tool calls it triggers; Turn = zero or more steps (opened when the first input is claimed, closed when nothing is owed).

The key extension points are all waterfalls: agent/pre-step (decides what the model sees, and can rewrite or even reject claimed messages), agent/request, llm/stream, tools/pre-execute|execute|post-execute. agent/turn-stopping is a serial terminal checkpoint. A rejected or empty first claim still closes a persistent turn that "spent zero steps", and the log records the attempt itself.

4. Capability Seams: Swap One Provider, Swap the Entire Product

A seam = a replaceable capability, with three roles: Service Definition (the declared interface) + Service Provider (the implementation) + Consumer (the user, usually a model-facing tool).

The example the documentation gives is telling: the filesystem and subprocess providers share the same execution world, and pointing them at a remote sandbox means Bash, PTY, and LSP all follow along, with no provider forking needed. The Subagent provider is likewise a broad spectrum behind one interface, from an in-process subagent to delegating a turn to another product.

From the auto-generated capability-seams documentation I counted: 40+ ctx services, covering llm / tools / sessions / fs / shell / subprocess / terminal / lsp / sandbox / approval / subagents / jobs / web / workflow / goals / skills / storage / telemetry / compaction / spill (overflow storage for oversized tool output) and more.

5. Scope: A Per-Agent World

Contributions (tools, prompt sections, variables, restrictions, listeners) are either global or scoped to a single scope key (by convention, a live agent is the key of its own scope). Shadowing: a scoped tool or section with the same name replaces the global one of that name, which is the mechanism behind per-agent personas and per-agent tool variants. Scope does not inherit to subagents; lineage (parentSession, delegationDepth) is data, not a visibility structure.


Review of Highlight Features

🔁 tool-cordis: the Agent modifies its own runtime. Five model-facing tools: cordis_inspect (a read-only report: services, plugin fibers, tools, dynamic packages), cordis_define (after syntax checking, records a new package, with a host half and a browser half), cordis_run (the host half evaluates in a vm sandbox, the browser half pushes to all open pages), cordis_stop / cordis_undefine. The official example examples/web-cordis/cordis.yml carries an exceptionally candid warning at the top:

Temporary Plugin code can reach every injected live capability; treat this deployment like shell access, not as a security boundary.

🎯 goal: a same-session objective domain. A durable completion objective attached to an existing session, with a revised state of active/paused/blocked/complete + a goal-round cap. Key design: goal activation is deliberately not persisted, so after a resume or fork, automated work cannot continue until it passes through a human-authorized /goal change.

🔁 Ralph loop: a fresh-agent workflow aimed at an immutable objective; each round is a brand-new sub-session (carrying no seed from the parent conversation), and cross-round state relies on the shared workspace plus a bounded structured handoff (status/summary/evidence/next steps/blocker).

🪝 hooks bridge: translates the hooks.json shell hook protocol of Claude Code and Codex into dsh's typed interception points, so your existing CC/Codex hook ecosystem can be mounted directly. Conversely, a "native hook" is just an ordinary Cordis plugin on these extension points.

🤖 subagent providers: in-process spawn / fork, ACP, Codex, Claude Code, dsh-sdk; delegation is an interface, not an implementation.

🔒 Sandbox: macOS uses sandbox-exec/seatbelt; Linux uses the in-house landlock-run (~300 lines of C11 against the raw Landlock kernel UAPI, statically linked with musl): self-restrict-then-exec, with the ruleset inherited across execve, so the caller is unrestricted while the entire called process family is fenced in, fail-closed (if the kernel cannot enforce it, execution is refused).

🐍 Python SDK: JSON-RPC over stdio driving the harness subprocess, with a high-level turns API; BENCHMARK.md points to using it for benchmarks.

Others: an E2B sandbox POC, an ACP automation server, OTel session telemetry, a worker-thread workflow engine, compaction (pre-step pressure + context overflow recovery, pruning tool results before summarizing), spill (oversized tool text spilled out + a model-visible locator).

21 model-facing tools: bash, bash-persistent, terminal, fs, fs-search, str-replace-editor, web, lsp, todo, goal, skill, subagent, subagent-control, subagent-report, jobs, session-query, ask-user, workflow, ralph, cordis, and more.


Mac Studio Field Test Log (Release Night)

1. Source installation: git clone --depth 1 + pnpm install took 18.7 seconds (pnpm 11.7.0). The host part of the source build passed; the client part errored at the TypeScript stage (Type 'bigint' is not assignable to type 'ReactNode', a known rough edge of the developer preview), so I bypassed it with the prebuilt npm package.

2. Web UI: npx @deepseek-ai/dsh web → dsh web: http://127.0.0.1:3080, HTTP 200, with the page bootstrapped by the __DSH_BOOT__ plugin graph (typert-registry → api-gateway → client-connection... the client is itself a plugin graph).

3. Model connection: The official DeepSeek API key returned Authentication Fails (governor) that night (authentication rate-limited or rejected). I switched to our own OpenAI-compatible gateway (Alibaba Cloud Bailian Token Plan) and, following the official documentation, wrote $DSH_HOME/settings.yaml:

agent-default-model:
  provider: bailian
  model: deepseek-v4-pro
llm-pi-ai:
  providers:
    bailian:
      displayName: Bailian Token Plan
      apiKeyEnv: BAILIAN_API_KEY
      api: openai-completions
      baseURL: https://token-plan.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1
      models:
        - id: deepseek-v4-pro
          contextWindow: 1000000
          maxTokens: 32768

apiKeyEnv is a credential reference; the key never enters the config file and is resolved on every request. A misconfigured reference fails explicitly with MISSING_CREDENTIAL rather than grabbing an unrelated key from the environment and muddling through. This "write-only credential + reference resolution" design deserves praise.

4. Headless end-to-end test:

npx @deepseek-ai/dsh --profile headless "Reply with exactly: HARNESS_ALIVE"
# EXIT=0, stdout: HARNESS_ALIVE, stderr empty

npx @deepseek-ai/dsh --profile headless "What model are you? Answer in one short sentence."
# EXIT=0, stdout: "I am a coding agent powered by DeepSeek v4 Pro (deepseek-v4-pro)."

End-to-end success: a one-shot session, prompt assembly, the model request, log persistence, printing the final answer to stdout, and a clean exit (turn/end completed → exit 0). Going from a custom gateway to DeepSeek V4 Pro required changing only one YAML file throughout, with no service restarts.


Comparison with Our Infrastructure

vs OpenClaw: OpenClaw's plugins are "channel + tool extensions", and its core loop is fixed; dsh's plugins are "the product itself", and even the loop is replaceable. OpenClaw's configuration is static JSON; dsh's is a layered patch system + runtime HMR. Their positioning differs: OpenClaw is a multi-channel AI assistant platform, and dsh is a developer-facing agent product foundation.

vs Prime Agent's Continual Harness (researched last week): Prime Agent's self-improvement happens at the text layer (the agent can CRUD its own prompts/skills/memory); dsh's tool-cordis self-modification happens at the code layer (defining and mounting new plugins at runtime, with the host half going into a vm sandbox and the browser half going into every page). Deeper, and more dangerous; even the official documentation says "treat like shell access".

vs our Loop Engineering: two designs resonate strongly with our practice,

  1. The "Model-visible ⟺ logged" invariant = the same conviction behind our external supervisor: the source of truth must live outside the LLM (our session JSONL scanning vs its SessionEvent log).
  2. The waterfall event chain = a typed implementation of our Gate concept: agent/pre-step can reject input, tools/pre-execute can intercept execution, and agent/turn-stopping is a terminal checkpoint.
  3. Non-persisted goal activation, with resume requiring human authorization = the same class of design as our "automated tasks must have a point where a human re-authorizes them".

Three things worth borrowing directly:

  • Docs as code: module-graph, capability-seams, tool-catalog, and config-catalog are all auto-generated by scripts from code declarations, with integrity guards (generated catalogs + a doc-sync gate). Our skill-system documentation can learn from this.
  • The session log invariant: anything model-visible must be reconstructible from the log. If the AK-SDD pipeline adopted this, it could fundamentally cure problems like "subagent-reported data being inconsistent with actual execution".
  • The three-role capability seam model: the separation of Service Definition / Provider / Consumer is more replaceable and testable than our current "tool = implementation" granularity.

Risks and Limitations

  1. Developer Preview: officially stated to include breaking changes; the SQLite schema and session format both declare "no compatibility promise". It is too early to put this into production now.
  2. Self-modification is an attack surface: tool-cordis's dynamic plugin code can reach every injected live capability, and the official documentation itself characterizes it as shell-access level.
  3. No official benchmark: BENCHMARK.md is just three lines on "how to run it", with no scorecard; capability claims currently rest entirely on documentation and demos.
  4. The model adapter is still narrow: natively there is only llm-deepseek + the generic llm-pi-ai (multi-provider); reasoning control and multimodal support are still being filled in.
  5. The community has just started: Discussions + Discord + a WeCom group, with zero issues (not yet open, or just opened).

Conclusion and Next Steps

DeepSeek Harness is the first agent harness to take the "plugin framework" to a paper-grade level of seriousness. Its bet is this: the long-term competitiveness of an agent product does not lie in any fixed loop design, but in composability itself: along the time dimension any component can be fully undone, and along the spatial dimension dependencies can be declared and made reactive. Add tool-cordis's self-modification capability, and this is a real path from slogan to engineering for a "self-evolving agent harness".

For us, the short-term strategy is to observe and borrow rather than migrate: watch for the first tagged release and the final version of the paper; borrow its generated-docs practice, its session log invariant, and its capability seam modeling. What is worth experimenting with in the medium term is the Python SDK + JSON-RPC to use dsh as the execution foundation for an AK-SDD-style pipeline, and the hooks bridge to reuse existing Claude Code hook assets.

Reference links:


This article is based on field testing on a Mac Studio in the early hours of 2026-08-14 (dsh 0.1.0-rc.6 + deepseek-v4-pro via our own gateway). The project is in a period of rapid iteration; details are subject to the official repository.