I’ve been saying for months that the model is not the product anymore, the harness is. Nobody believed me until I sat down and put three of them side by side on the same machine, and what I found made me rewrite this piece twice.
Here’s the thing that started this whole rabbit hole: take the exact same underlying model, no prompt changes, no fine-tuning, no new training data, just swap the execution environment around it, and the model goes from mediocre to genuinely good. I’d seen this claimed before in passing, but didn’t take it seriously until I found the number that made me stop scrolling. Research cited by JIN, who writes the System Architect newsletter on Medium, put a figure on it: the same model, wrapped in a different harness, moved a programming benchmark from 42 percent success to 78 percent. No new weights. Just a different thing deciding how context gets assembled, how tools get called, and how much rope the model gets before someone steps in.
That’s not a small effect. That’s the difference between “would not ship this” and “this is genuinely useful,” and it comes entirely from the layer wrapped around the model.
So I spent a couple of weeks doing what most comparison posts skip: installing the harnesses, reading the code that actually ships, not just the README, and finding where each one’s design choices show up as real, felt friction. Here’s what I found with DeepSeek Harness, OpenAI’s Codex as a Platform, and, as the baseline I use every day, Claude Code.
DeepSeek Harness: everything really is a plugin
DeepSeek Harness launched as a developer preview a few months ago under the tagline “Everything is a Plugin,” and the GitHub numbers aren’t exaggerated marketing, the repo is sitting north of 180,000 stars with just under 20,000 forks. That’s a real, fast-growing community. It’s MIT licensed and fully open source.
The architecture is the interesting part, not the star count. DeepSeek Harness is built on Cordis, a plugin meta-framework the team vendored and patched rather than pulled in as a stock dependency. Cordis treats every capability, the agent loop, the tool scheduler, the session log, the sandbox, the web UI, even the SDK itself, as a plugin implementing a Service interface. Plugins declare what they need through an inject field, and Cordis resolves load order from those dependencies instead of you hand-writing a boot sequence. A context object works as a shared registry where services claim stable keys like ctx.tools, ctx.llm, or ctx.sessions, so a plugin asks the context for whatever holds that key rather than importing a concrete implementation directly. Registrations install through ctx.effect(), which makes them reversible, so a reload or teardown cleanly unwinds instead of leaving dangling state behind. That's a genuinely clean design for hot-swapping components, and it's why you can, in principle, swap the model provider, the tool set, or the entire workflow engine without touching the rest of the system.
On paper this is exactly what “everything is a plugin” should mean. In practice, I wanted to know what actually ships versus what the pitch implies, so I read through the repository the way you’d read a codebase you were about to depend on in production, not the way you read a landing page.
What’s actually in the box
The core repo runs to roughly 453,000 lines of TypeScript across around 219 workspace packages, built on a vendored fork of Cordis pinned at 4.0.0-rc.7 with eighteen local patches on top of it, shipping as npm package 0.1.0-rc.6. That's a lot of surface area for something billed as "just plugins": you're depending on a patched fork of a pre-1.0 framework, not the framework's own release train.
The default model catalog is narrow, exactly two models, deepseek-v4-flash and deepseek-v4-pro. Everything else, Anthropic, OpenAI, Bedrock, Vertex, Azure, or any custom OpenAI-compatible endpoint, gets wired up through the provider configuration screen, not auto-discovered. This is where the interesting part starts, because a fair amount of what happens once you add providers isn't obvious from the docs alone.
The install, the router, and the two files nobody reads twice
I leaned heavily here on a hands-on writeup by a developer who goes by Andrus, published on the AI Advances gopubby publication, titled “Should You Switch From Claude Code to DeepSeek Harness? I Read the Shipped Code Instead of the README.” The framing matches exactly what I found poking around myself: read what ships, not what the marketing implies.
First number that jumps out: the install is 359 megabytes before you type a single character into it. That’s the @deepseek-ai/dsh package on disk, before any plugins, before any provider config. That's roughly the size of a small Electron app, for a CLI-first agent harness.
Second, and this is the part that made me go dig through the provider code myself: what looks in the repo like a small test fixture for provider configuration is, once you trace how it’s actually wired at runtime, a full multi-provider router covering somewhere in the neighborhood of three dozen providers, Anthropic and Claude models included. That’s not disclosed prominently in the getting-started docs, you find it by reading docs/user/guide/providers.md and following the "Add provider" flow into the actual adapter code. It's not sinister, DeepSeek genuinely wants to let you run any model through it, but the docs sell "DeepSeek's own harness" and the shipped code is a general-purpose model router that happens to default to DeepSeek's models. Worth knowing if procurement ever asks whether this sends anything to a competitor's API: the honest answer is it can, and the plumbing is already built in, quietly.
Third, and the one that surprised me most: there is no native local-model provider. Not Ollama, not LM Studio, nothing with a first-class entry in the provider catalog. Andrus was running a 36GB MacBook Pro with Ollama serving a locally quantized Qwen model handling roughly 70 percent of his day-to-day workload, and DeepSeek Harness could not drive that model without him hand-writing a custom provider entry pointed at Ollama’s OpenAI-compatible endpoint. Compare that to Codex, which documents OSS mode against Ollama or LM Studio directly in its own config guide. DeepSeek Harness technically supports it, since any endpoint speaking the OpenAI chat completions shape can be registered, but “supports it if you write the adapter yourself” is a different bar than “ships with it.”
Fourth is one I verified myself against the live repository rather than repeat secondhand. DeepSeek Harness’s own repo carries both a CLAUDE.md and an AGENTS.md at the root, a common pattern now that multiple coding agents look for different filenames. I pulled both files directly. In the dsh repo itself, CLAUDE.md is nine bytes, literally just the text AGENTS.md, a pointer telling any tool that reads it to go read the real file instead: the maintainers being disciplined about not duplicating content across two files.
The problem is that discipline is a convention, not something the harness enforces for you. Point dsh at a typical real-world repo where CLAUDE.md and AGENTS.md have organically diverged, which happens constantly once a team runs Claude Code and a second agent side by side for a while, and the context loader pulls both files into every prompt, in full, with no deduplication. I ran a rough estimate against two files from one of my own projects: a CLAUDE.md around 1,400 tokens of Claude Code-specific tool conventions, and an AGENTS.md around 2,100 tokens of generic instructions that had drifted out of sync with it over eight months. That's an extra 1,400 tokens of pure duplication tax on every single turn, not once per session, because there's no caching discount for instruction files re-sent with each prompt. Over a long session with fifty or sixty tool-calling turns, that's tens of thousands of wasted tokens paying for context the model already saw.
None of these four things is disqualifying on its own. Together they’re a pattern: DeepSeek Harness’s actual behavior, once you read the shipped code, is broader and less disciplined than the “everything is a plugin, clean and modular” pitch implies. The plugin architecture underneath is real and well-engineered. What ships around it needs a closer read than the README gives you credit for needing.
Codex as a Platform: OpenAI’s bet on the app-server, not the chat window
OpenAI’s framing for Codex shifted meaningfully this year. It’s no longer positioned as a coding chatbot with a CLI, it’s positioned explicitly as an open agent harness that other products build on top of, laid out in their “Codex as a Platform” writeup on the OpenAI Developers site.
The architecture splits into three integration surfaces instead of one monolithic app. codex exec is for scripts and bounded, one-off background tasks, the kind of thing you'd trigger from a CI job. The Codex SDK is a direct programmatic interface for applications that need to start, resume, or stream a task from inside their own code. The app-server is the deepest integration point: it exposes persistent conversation state, event streaming, tool exposure, and approval-request handling, so a product can build its own UI around Codex's agent loop while Codex handles sandboxing, execution, and approval plumbing underneath. The dividing line is deliberate: your application owns business logic, data sources, MCP services, approval policy, and the UI. Codex owns the agent loop and the sandboxed execution environment. That's a service boundary, not a chat window you embed.
The results backing this up were more concrete than I expected. On ARC-AGI-3, retained reasoning combined with context compaction pushed GPT-5.6’s accuracy from 13.3 percent to 38.3 percent, while cutting output tokens roughly sixfold. Same shape of claim as the 42-to-78 harness effect I opened with: same reasoning capability, dramatically different result once the surrounding execution logic changes how context gets carried and compacted.
The case study that actually convinced me this isn’t just a demo-day number is Asana’s public writeup of migrating their frontend test suite off Enzyme, an aging testing framework that had become a hard blocker to any React upgrade, onto React Testing Library. Asana’s engineering team estimated the migration properly done, framework swap plus improved coverage plus cleaning up legacy test infrastructure, would take roughly five more years at their existing pace. They ran it with Codex instead, frontier models on a high-reasoning setting, up to four agents running in parallel across different directories, working through nights and weekends unsupervised. The prompt driving all of it was five sentences: migrate everything using Enzyme to React Testing Library, follow existing conventions, test the changes, bias toward easier conversions first. It shipped in about a week and a half of engineering time spread across two calendar weeks. Total cost, model usage plus infrastructure, landed around 12,000 dollars, against a back-of-envelope estimate of roughly 6 million dollars for the manual version of the same scope.
Asana’s postmortem is refreshingly honest about why it worked, and it’s not “the model is smart.” Their conclusion was that agent output quality depends heavily on the quality of the environment you hand it: established conventions, clear examples, well-designed test helpers already in the codebase before any agent touched it. Their friction points weren’t model limitations either, they were stale internal documentation steering agents toward deprecated patterns, and infrastructure bottlenecks like slow linting and flaky tooling creating more drag than the model ever did. That’s a useful data point about where the ceiling actually sits on this kind of work: not model capability, environment quality.
If you want Codex against a local model rather than a hosted one, it supports this natively through OSS mode, unlike DeepSeek Harness’s roll-your-own approach:
# ~/.codex/config.toml
[model_providers.ollama]
name = "Ollama (local)"
base_url = "http://localhost:11434/v1"
wire_api = "chat"
[profiles.local]
model = "qwen2.5-coder:32b"
model_provider = "ollama"
# pull a coding-capable model locally first
ollama pull qwen2.5-coder:32b
# then run codex against your local profile, no API key needed
codex --profile local
That’s a first-class, documented path, not a workaround you have to reverse-engineer from the adapter source.
The sharpest framing I found: it’s not features, it’s where complexity accumulates
The comparison that changed how I think about this space didn’t come from a feature checklist. It came from JIN’s argument that the reason everyone is suddenly rewriting their harness runtime, instead of just picking a bigger model, is that the harness is now a first-class engineering decision. Once you accept that, the interesting question stops being “which harness has more plugins” and becomes “where does each harness let complexity pool up, and is that where you actually want it.”
DeepSeek Harness pools complexity into the plugin graph. The core, Cordis plus the agent loop, is genuinely compact and well-factored, but almost everything you’d care about (which provider you’re really talking to, how instruction files get merged, what a plugin does at runtime versus what its name implies) lives inside a plugin you have to go read. That’s a compact-core design: power and risk both concentrated in a growing, loosely-typed plugin ecosystem around a small, disciplined center.
Codex pools complexity at the boundary between your application and the agent loop. The app-server model means Codex stays relatively opinion-free about approvals, business logic, and UI, and pushes that decision-making onto whoever builds on top of it. That’s a service-boundary design: the complexity doesn’t disappear, it moves to whichever side of the API you’re standing on, and if you’re the one building the product around Codex, that’s you.
Claude Code, by contrast, keeps a documented event surface, session lifecycle, per-turn hooks, PreCompact and PostCompact around context management, tool-use hooks with allow-or-deny authority, that behaves like a replayable, inspectable runtime rather than either a plugin soup or a bare API boundary. You don’t get Cordis-style hot-swapping of the core loop, and you don’t get Codex’s clean hand-off to an external app, you get a system that logs its own decisions at each lifecycle point and gives you a hook into every one, a different bet: less raw extensibility, more auditability of what already happened.
None of these is objectively correct. They’re three answers to the same question of where you want the hard-to-reason-about parts of an agent system to physically live.
I wanted numbers behind this framing rather than vibes, so I went looking for a benchmark that measures harnesses as a variable rather than models. The closest real one is Harness-Bench, a 2026 paper measuring harness effects across models on 106 realistic agent tasks. Worth being upfront about its limits before citing numbers from it: it tests six configurable, generically-named harness configurations plus Codex evaluated separately, and does not test DeepSeek Harness or Claude Code by name, so treat this as evidence that harness choice matters at all, not a scoreboard ranking these three products. With that caveat attached, the spread was large: the best configurable harness hit a 76.2 completion score at 68,700 tokens per task, the weakest scored 52.4 at 82,100 tokens, a 23.8-point gap while burning more tokens for a worse result. Codex, evaluated independently, scored 80.4. The authors are explicit their scoring relies partly on LLM-assisted judging and describe the results as diagnostic under one fixed protocol, not a guarantee of production performance. I’d treat it the same way: directionally convincing that harness architecture is a real, measurable variable, not a leaderboard to build a purchasing decision on alone.
A separate paper, an enterprise agent evaluation framework called CLEAR, found something that rhymes with this from a different angle: across six agent systems on 300 enterprise tasks, up to 50x cost variation showed up between systems delivering similar accuracy, with accuracy-optimized configurations running 4.4 to 10.8 times more expensive than cost-aware ones for comparable quality. Different study, same underlying point: the harness and its configuration choices do as much work on your bill as the model does.
The comparison, side by side
+----------------+-----------------------+-----------------------+-----------------------+
| Dimension | DeepSeek Harness (dsh) | Codex as a Platform | Claude Code |
+----------------+-----------------------+-----------------------+-----------------------+
| Local models | No native provider, | Native OSS mode via | None. Model call is |
| | hand-wire a custom | Ollama/LM Studio, | inherently Anthropic |
| | OpenAI-compat endpoint | documented in config | API, Bedrock, or Vertex|
+----------------+-----------------------+-----------------------+-----------------------+
| Plugin system | Core design principle, | Not plugin-based, app | Hooks + MCP + |
| | Cordis DI, ctx keys, | owns logic, Codex owns | subagents, fixed core |
| | reversible effects | the loop via app-server | with attach points |
+----------------+-----------------------+-----------------------+-----------------------+
| Context handling| Loads CLAUDE.md and | App controls context | CLAUDE.md at session |
| | AGENTS.md in full, no | assembly, Codex handles | start, PreCompact/ |
| | dedup if they diverge | compaction internally | PostCompact hooks |
+----------------+-----------------------+-----------------------+-----------------------+
| Rough cost | Cheap model defaults, | Frontier pricing, but | Frontier pricing, |
| | but router overhead and | Asana case shows ~500x | subscription tiers plus |
| | duplicate context add | cheaper than the manual | usage-based options |
| | unseen tax | alternative | |
+----------------+-----------------------+-----------------------+-----------------------+
That cost row deserves its own caveat: “cheap model defaults” and “frontier model cost” aren’t apples to apples, DeepSeek Harness defaults to DeepSeek’s own inexpensive-per-token models, while Codex and Claude Code default to frontier-tier pricing. The point isn’t that one is cheaper in isolation, it’s that DeepSeek Harness’s per-token savings can get quietly eaten by the router overhead and the duplicate-context tax described above, in a way that doesn’t show up until you actually watch token usage over a real session instead of trusting the sticker price.
Where I actually landed
If you’re a solo developer or small team that wants to swap models freely, including cheap providers or your own fine-tunes, and you’re comfortable reading plugin source when the docs run thin, DeepSeek Harness’s architecture is legitimately good engineering. Just budget time to read the provider code before trusting what’s actually being called, and check your CLAUDE.md/AGENTS.md situation before assuming the context loader is economical with your tokens, because right now it isn't.
If you’re building a product that needs an agent embedded inside it, not a chat window bolted on, Codex’s app-server model is the more honest architecture, specifically because it draws the boundary explicitly instead of leaving you to reverse-engineer where your responsibility starts. The Asana numbers are extraordinary, but the caveat that matters is Asana’s own: none of it works without a codebase that already has clean conventions and good test scaffolding for the agent to imitate. Codex won’t invent that discipline for you, it amplifies whatever discipline is already there, in both directions.
If you’re already inside the Claude Code ecosystem and mostly want a documented, inspectable event surface to build policy on top of, the hooks-plus-sandbox model is still the one I reach for by default, mainly because auditability of what happened beats raw extensibility of what could happen, for work where an agent runs semi-unsupervised against a codebase people’s paychecks depend on.
None of these are wrong choices. They’re three bets about where complexity should live, and the honest answer is the same one that’s true of any infrastructure decision: read the code that actually ships, not the page trying to sell it to you.
Tags: agent-harness, deepseek, openai-codex, claude-code, ai-agents, developer-tools, llm-engineering
Top comments (0)