DeepSeek Harness is an open-source runtime for AI agents. DeepSeek labels the project a developer preview, so builders should expect breaking changes.
DeepSeek built Harness as an open-source runtime for AI agents. The project uses one formula:
Agent = Model + Harness
The model handles reasoning. The harness provides tools, memory, sessions, permissions, and an interface. DeepSeek lets developers replace and recombine those parts through plugins.
This article introduces DeepSeek Harness and breaks down how it works: its plugin architecture, operating modes, session model, and plugin workflow.
A short walkthrough:
What shipped
Repo: deepseek-ai/deepseek-harness
DeepSeek released it on 13 Aug 2026 under the MIT license. The TypeScript monorepo remains in developer preview, so expect compatibility-breaking changes.
Quick start:
npx @deepseek-ai/dsh web
Web UI defaults to http://127.0.0.1:3080. Add an API key under Settings -> Models. Pick a workspace before the input box unlocks.
From source:
git clone https://github.com/deepseek-ai/deepseek-harness.git
cd deepseek-harness
pnpm install
pnpm run build
pnpm dsh web
Inspect the composed config without booting:
dsh --profile web --dump-config
Headless one-shot:
dsh --profile headless "Summarize this repo and print the result"
The claim that matters
Most coding agents still have a privileged core. You can hang skills, hooks, MCP, or plugins off it. You cannot replace the loop, the session log, or the tool pipeline without forking the product.
DeepSeek's architecture doc is blunt: there is no privileged core to patch. You extend dsh by mounting a plugin next to the others. Registrations are effects. Unload the plugin and they reverse.
That is why "everything is a plugin" is stronger than "we have an extension API."
The slogan depends on three details: plugins can be replaced, their effects can be reversed, and the runtime keeps session state consistent.
1. Registration is an effect. Unload rolls it back.
If a plugin can own the loop or the session store, leftover listeners after a hot-reload are not a cosmetic bug. They are a second loop.
Official wording: registrations are effects that unwind when their plugin unloads. That is why they dare to expose UI, storage, and the loop as plugins. You can swap the dangerous pieces without leaving half an old implementation in the process.
Replaceable is the slogan. Reversible is the capability.
2. Anything the model saw must be findable in the log.
The session is not a chat transcript. It is an append-only event log. History for the next request is derived from that log. Stream chunks go back into the same place.
The log includes request metadata such as the provider, model, system prompt, and tool catalog. A resumed session can reconstruct what the model saw before the interruption.
Official invariant is the same idea: if the model saw it, the log can rebuild it. Compress, truncate, or summarize without a path back, and the resumed agent is a different agent.
One log, five jobs: resume, fork, replay, search, audit.
3. Tools may finish out of order. Write-back stays in model order.
This one is small and easy to miss. Official tools README: the loop groups consecutive concurrency-safe calls into a bounded rolling pool. Exclusive calls are ordering barriers. Only dispatch and body overlap. Policy, durable results, and context keep model order.
Code Mode reuses the same contract. Sub-calls start in submission order. maxParallelSubCalls defaults to 10. Set it to 1 and you are serial again.
If a fast search jumps ahead of a slow file edit in the context, the causal chain the model planned no longer matches what it sees. That will not throw. Later reasoning just quietly goes wrong.
Cancel has the same taste: started work tries to settle, work that never started still gets a synthetic result. The log should not be missing half a tool chain.
The three details point at one test: a runtime is reliable after it is interrupted, replaced, or cancelled, not merely after it boots.
The kernel is Cordis. The paper behind it, A Programming Paradigm for Spatiotemporal Composability, names two properties:
- Time: uninstall should undo the plugin's side effects.
- Space: plugins declare dependencies and react when the graph changes.
In practice you write a TypeScript module that exports apply(ctx):
import type { Context } from '@deepseek-ai/cordis'
export const name = 'hello-plugin'
export function apply(ctx: Context) {
console.log('[hello-plugin] plugin loaded!')
}
Need tools? Declare inject = ['tools'] and register through ctx.tools. The loader waits until the service exists. You do not hand-order startup.
A running dsh is a plugin tree stacked in order:
- bundles listed by the profile (
dsh-base, then extras) - the profile's
cordis.patch.yml - home-level
$DSH_HOME/cordis.patch.yml - any
--patchoverlay
A patch finds a row by id and replaces the whole config. It does not deep-merge keys. First-time users lose API keys this way.
I counted 219 workspace packages under packages/*/*. That is the cost of the breadboard: every seam (interface, provider, consumer) can live in its own package.
Four modes, four jobs
Official site lists four presets:
| Mode | What it is | Use it when |
|---|---|---|
| Standard | Full coding agent: files, shell, search, skills, plan, goals, subagents, workflows | Daily work |
| Code | Standard plus Code Mode SDK. Model writes one TypeScript program that batches tool calls | Long, repetitive tool pipelines |
| Minimal | Persistent bash + str_replace_editor |
Model benchmarks, thin evals |
| Creator | Standard plus runtime inspection and in-memory plugin experiments | Authoring presets |
Code mode is programmatic tool calling (PTC). It is not Claude Code Dynamic Workflows. PTC compresses one agent's tool loop into a program. Dynamic Workflows fan a large task out across many agents. Different layer. Official Code Runtime today: TypeScript only, worker-thread backend. The isolation field is explicitly not a security claim.
Workflows can fan out subagents from a model-written script. Official limits: foreground only, no checkpoint/resume, no nested workflows.
Subagent providers in the capability map already include in-process spawn/fork, ACP, Codex, and Claude Code. The transport is swappable. The contract stays.
Why the session log is the real product
DSH stores an append-only SessionEvent log. Official invariant: if the model saw it, the log can rebuild it. Runtime asserts this.
That is why Trajectory view exists. Resume, fork, search, replay, and telemetry all project from the same stream. Codex has history. DSH treats the log as architecture.
Sandbox vocabulary already converged with Codex: read-only, workspace-write, danger-full-access. Official README: this vocabulary is file effects only. No network, process, or syscall policy. No usable backend means SANDBOX_UNAVAILABLE. Fail closed. Policy follows the call, not the provider.
What the community is building
Capability extensions
ModLens gives text-only models vision by turning pasted images into structured evidence. Graph Memory carries useful context across sessions. Browser plugins let the agent work with websites, while dsh-github-connector can create, review, and merge pull requests from the conversation.
New ways to work
dsh-TUI replaces the Web UI with a terminal interface. dsh-agent-teams lets one agent coordinate a team of agents. DSH Better Sidebar adds file editing, terminal, Git, and subagent panels to the browser.
Personal interfaces
Petdex adds animated desktop pets that react to agent activity. dsh-minigames puts offline games in the sidebar. dsh-ads turns the interface into a parody of a 2005 web portal.
Community developers use one extension system to reshape the agent's capabilities, workflow, and interface.
How to install a plugin without guessing
Local overlay for learning:
pnpm dsh web --patch ./scratch-plugin/cordis.yml
Real distribution is a bundle. package.json declares dsh.bundle. Users install into a profile:
dsh plugin --profile demo add ./hello-plugin
dsh plugin --profile demo add github:you/hello-plugin
Git installs pull source, not build artifacts. Each package needs a self-contained prepare script. pnpm >=10 refuses that script until the user allows it in pnpm-workspace.yaml. That code runs on your machine, outside the agent sandbox. Lock the commit SHA.
A third-party plugin in your profile is host code. It runs outside the model sandbox with permissions close to the dsh process. Review its install script, dependencies, and host code as you would any local program.
Start with workspace-write plus per-call approval. Before installing a third-party plugin, have a trusted model review the code.
Python path exists too: pip install deepseek-harness-sdk. The shipped minimal combo is bash + editor, danger-full-access, Linux or modern macOS arm64. Isolated workspace only.
Who should use it now
DeepSeek still labels Harness a developer preview. Expect its configuration and APIs to change.
It fits:
- teams building agent infrastructure
- engineers studying agent loops, sessions, and plugin architecture
- teams that need control over models, tools, storage, or UI
Wait for a later release if:
- you want a stable coding agent that works out of the box
- you do not want to manage configuration, compatibility changes, or plugin security
Use it for study and experiments. Treat production adoption with more caution.

