Somewhere in the last year, “harness” quietly became the most overloaded word in agentic AI. I’ve seen it used to mean a system prompt. I’ve seen it used to mean an entire CLI product with a sandbox, a permission model, and a plugin marketplace. I’ve seen a Twitter thread call a single evaluator loop a harness, and a week later watched someone else use the exact same word to describe a 450,000-line TypeScript codebase with 219 packages. Everyone nods along like we’re all talking about the same thing. We are not.
So I spent about a week doing the thing that actually resolves this kind of confusion: I stopped reading vendor pitches and started building. I wired up a taxonomy against the actual academic papers it traces back to, wrote runnable code for the five layers that separate a real harness from a system prompt with delusions of grandeur, installed three open-source harnesses that disagree with each other on almost every design decision, and spent a long afternoon with Claude Code’s built-in graph runtime trying to figure out when a loop actually deserves to become a graph instead of just looking like one on a whiteboard.
This is the writeup. It has more code than opinions, and where I do have opinions I’ve tried to mark clearly which claims are load-bearing and proven versus which ones are still a single vendor’s benchmark or an unreplicated arXiv result.
A real taxonomy, not a vibe
Before any of the layers or the code, it’s worth pinning down what we’re actually taxonomizing, because “agent” and “harness” get used interchangeably in ways that hide a genuinely useful distinction.
Anthropic’s own field guide on building effective agents, published by their applied AI team, draws a line I now use as a first filter for every system I design: workflows are systems where an LLM and its tools are orchestrated through predefined code paths, and agents are systems where the LLM dynamically directs its own process and tool usage, maintaining control over how it accomplishes a task. Both are legitimate. The mistake is reaching for the second when the first would be cheaper, faster to test, and easier to hand to a teammate who has to maintain it after you.
Here’s the map I ended up with, tied back to where each pattern actually comes from instead of a marketing deck:
+----+------------------------+---------------------------------------------------+
| # | Pattern | Origin / grounding |
+----+------------------------+---------------------------------------------------+
| 1 | Prompt Chaining | Anthropic, Building Effective Agents (2024): |
| | | fixed sequence of LLM calls, each output feeds the |
| | | next, with a programmatic gate in between |
+----+------------------------+---------------------------------------------------+
| 2 | Routing | Anthropic, Building Effective Agents: classify the |
| | | input, send it down one of several fixed paths |
+----+------------------------+---------------------------------------------------+
| 3 | Parallelization | Anthropic, Building Effective Agents: sectioning |
| | | (independent subtasks) or voting (same task, N runs) |
+----+------------------------+---------------------------------------------------+
| 4 | Orchestrator-Workers | Anthropic, Building Effective Agents: a central LLM |
| | | decomposes a task dynamically and dispatches workers |
+----+------------------------+---------------------------------------------------+
| 5 | Evaluator-Optimizer | Anthropic, Building Effective Agents: one LLM |
| | | generates, a second critiques, loop until it passes |
+----+------------------------+---------------------------------------------------+
| 6 | ReAct | Yao et al., "ReAct: Synergizing Reasoning and |
| | | Acting in Language Models," arXiv:2210.03629 (2022): |
| | | interleave thought, action, observation, repeat |
+----+------------------------+---------------------------------------------------+
| 7 | Reflexion | Shinn et al., "Reflexion: Language Agents with |
| | | Verbal Reinforcement Learning," arXiv:2303.11366 |
| | | (2023): a failed attempt gets a verbal self-critique |
| | | that persists into the next episode as memory |
+----+------------------------+---------------------------------------------------+
| 8 | Graph Orchestration | Emergent 2025-2026 practice, formalized by tools like |
| | | LangGraph and, as of this year, Claude Code's own |
| | | dynamic workflows: control flow as an explicit, |
| | | inspectable graph instead of one agent's turn-by-turn |
| | | judgment |
+----+------------------------+---------------------------------------------------+
| 9 | Swarm | Peer agents coordinate directly with no central |
| | | supervisor. Powerful for debate and adversarial |
| | | red-teaming, hardest of the nine to keep bounded |
| | | because nobody owns the stop condition |
+----+------------------------+---------------------------------------------------+
Two things jumped out once I laid this out side by side instead of reading about each pattern in isolation. First, ReAct and Reflexion aren’t competitors to the five Anthropic workflow patterns, they’re a different axis entirely: the five workflow patterns describe how control flow is shaped, while ReAct and Reflexion describe what happens inside a single node once it starts reasoning. You can run a ReAct loop inside an orchestrator-workers worker. I do, constantly. Second, graph orchestration and swarm are really just the two ends of a spectrum of how much you’re willing to let the model itself decide the topology at runtime, with graph orchestration keeping a human-authored (or, as I found out later, Claude-authored) script in charge, and swarm handing that authority to the agents themselves.
I’ll come back to graph orchestration in detail near the end, because it’s the pattern that changed the most for me this year and it deserves its own section with real numbers instead of a table cell.
The five layers that make a harness a harness, not a prompt
Here’s the sentence that reoriented how I think about this whole space, and I’m stealing it because it’s more precise than anything I’d have written myself: your agent is six lines of Python, the harness is everything else. Paolo Perrone’s writeup on agent harnesses breaks that “everything else” into five layers, and I rebuilt each one myself to see which parts were genuinely load-bearing versus decorative. Short version: all five are load-bearing, and skipping any one of them shows up as a specific, predictable failure mode within about a day of real use.
Layer one: the execution boundary
The execution boundary is the code that runs between the model deciding to call a tool and the tool actually executing. If that boundary doesn’t exist, your only defense against a bad tool call is asking the model nicely in the system prompt not to do the bad thing, and I can tell you from direct experience that this fails. I tried telling a coding agent “never run destructive git commands without asking” in the prompt eleven separate times across eleven separate wording attempts before I gave up and wrote an actual gate.
# execution_boundary.py
# A minimal pre-dispatch hook you can drop in front of any tool-calling loop.
# It runs BEFORE the tool executes, and it can deny the call outright.
import re
from dataclasses import dataclass
@dataclass
class ToolCall:
name: str
arguments: dict
class ExecutionBoundary:
def __init__ (self):
self.deny_patterns = [
re.compile(r"rm\s+-rf\s+/(?!\S)"), # rm -rf / with nothing else appended
re.compile(r"git\s+push\s+.*--force"),
re.compile(r"DROP\s+TABLE", re.IGNORECASE),
]
def check(self, call: ToolCall) -> tuple[bool, str]:
if call.name == "bash":
command = call.arguments.get("command", "")
for pattern in self.deny_patterns:
if pattern.search(command):
return False, f"blocked by policy: matched {pattern.pattern}"
if call.name == "write_file":
path = call.arguments.get("path", "")
if path.startswith("/etc") or path.startswith("/sys"):
return False, f"blocked by policy: writes outside project scope"
return True, "allowed"
# usage inside your dispatch loop
boundary = ExecutionBoundary()
def dispatch(call: ToolCall, tool_registry: dict):
allowed, reason = boundary.check(call)
if not allowed:
return {"error": reason, "blocked": True}
return tool_registry[call.name](**call.arguments)
The thing I want to be honest about here: this is a denylist, and denylists are always incomplete. The real production version of this pattern, the one Claude Code ships as PreToolUse hooks, is not meaningfully more sophisticated in shape, it's the same before-the-call interception point, just with a richer set of things you can inspect and a JSON contract for the deny decision. The boundary doesn't make the model smarter. It just means the dumbest possible mistake gets caught in code instead of in production.
Layer two: sandboxing
This is the layer where I want to be the most careful about proven versus claimed, because the marketing in this space is aggressive and the underlying infrastructure genuinely is good, which makes it easy to blur the two.
What’s proven: container isolation via Linux namespaces and cgroups is decades-old, well understood, and fast, with startup in the low milliseconds. What’s also proven, not a vendor claim, is that a shared host kernel is the actual attack surface. A kernel bug or a container escape misconfiguration gives an attacker the host, and this has happened in the wild more than once. That’s why untrusted or model-generated code shouldn’t run in a bare container next to anything you care about.
The next tier up is microVMs, and Firecracker is the reference implementation, originally built at AWS for Lambda and Fargate. What’s proven there, from AWS’s own published numbers and reproduced independently since: each microVM gets its own kernel, fully separated from the host, with roughly 125 milliseconds of startup time, under 5 MiB of memory overhead per VM, and density up to around 150 VMs per second on a single host. Those are real, measured numbers from a system running in production at enormous scale, not a benchmark from a launch blog post.
What’s still closer to a claim I’d flag rather than repeat as settled fact: any specific vendor’s assertion that their hosted sandbox product is faster or cheaper than a competitor’s, when that number comes from the vendor’s own comparison page rather than a reproducible third-party benchmark. I’m not going to name a winner there, because I haven’t independently reproduced any of those specific throughput numbers myself, and I’d rather tell you what I didn’t verify than dress it up as more certain than it is.
Here’s the sandboxing pattern I actually run locally, using Docker as the accessible entry point, with the isolation flags that matter most and a clear note on where a container’s guarantees end and a microVM’s begin:
# sandbox_runner.py
# Runs untrusted, model-generated code in an isolated container.
# Requires Docker running locally: no cloud dependency, no API key.
import subprocess
import tempfile
import os
def run_sandboxed(code: str, timeout_seconds: int = 10) -> dict:
with tempfile.TemporaryDirectory() as tmp:
script_path = os.path.join(tmp, "snippet.py")
with open(script_path, "w") as f:
f.write(code)
result = subprocess.run(
[
"docker", "run", "--rm",
"--network", "none", # no outbound network at all
"--memory", "256m",
"--cpus", "0.5",
"--read-only", # root filesystem is read-only
"--tmpfs", "/tmp:size=64m", # writable scratch space only
"--security-opt", "no-new-privileges",
"--cap-drop", "ALL",
"-v", f"{tmp}:/sandbox:ro",
"python:3.12-slim",
"python", "/sandbox/snippet.py",
],
capture_output=True,
text=True,
timeout=timeout_seconds,
)
return {
"stdout": result.stdout,
"stderr": result.stderr,
"exit_code": result.returncode,
}
if __name__ == " __main__":
output = run_sandboxed("print(2 + 2)\nimport os\nprint(os.listdir('/'))")
print(output)
That --network none flag is doing more real security work than anything else in the command, and it's the one I see people skip most often because it breaks the convenient case of "let the agent pip install something mid-run." If the code genuinely needs network access, that's a decision to make explicitly, on a specific allowlisted domain, not a default. For the step up from this container to actual microVM isolation without running your own Firecracker fleet, Alibaba's OpenSandbox (Apache 2.0, on GitHub as alibaba/OpenSandbox) gives you a unified API in front of Docker, Kata Containers, gVisor, or Firecracker depending on what you configure, which is the honest self-hosted answer to "I want E2B's isolation model without E2B's bill."
Layer three: memory persistence
Short-term memory, the running transcript, is not the layer that gets people in trouble. It’s persistence across sessions, and specifically the fact that entries in a persistent store don’t expire unless something explicitly expires them. I’ve had an agent confidently cite a decision that had been reversed three weeks earlier because nothing ever told the store the old entry was stale.
Here’s a minimal, dependency-free persistence layer using SQLite, which is the honest answer for anyone who reaches for a vector database before they’ve proven they need one:
# memory_store.py
# Persistent, file-backed memory with explicit expiry. No external services.
import sqlite3
import time
import json
class MemoryStore:
def __init__ (self, path: str = "agent_memory.db"):
self.conn = sqlite3.connect(path)
self.conn.execute("""
CREATE TABLE IF NOT EXISTS memories (
key TEXT PRIMARY KEY,
value TEXT NOT NULL,
created_at REAL NOT NULL,
expires_at REAL
)
""")
self.conn.commit()
def remember(self, key: str, value: dict, ttl_seconds: float | None = None):
expires_at = time.time() + ttl_seconds if ttl_seconds else None
self.conn.execute(
"REPLACE INTO memories (key, value, created_at, expires_at) VALUES (?, ?, ?, ?)",
(key, json.dumps(value), time.time(), expires_at),
)
self.conn.commit()
def recall(self, key: str) -> dict | None:
row = self.conn.execute(
"SELECT value, expires_at FROM memories WHERE key = ?", (key,)
).fetchone()
if row is None:
return None
value, expires_at = row
if expires_at is not None and time.time() > expires_at:
self.conn.execute("DELETE FROM memories WHERE key = ?", (key,))
self.conn.commit()
return None
return json.loads(value)
if __name__ == " __main__":
store = MemoryStore()
store.remember("project:api-style", {"framework": "FastAPI", "auth": "JWT"}, ttl_seconds=86400 * 30)
print(store.recall("project:api-style"))
For anything past a few thousand entries where semantic lookup actually matters, swap the table for a local vector index like sqlite-vec or Chroma running on your own disk before you reach for a hosted vector database. The expiry column is the part that actually matters, not the storage backend.
Layer four: verification loops
This is the evaluator-optimizer pattern from the taxonomy above, wired as a harness layer with a hard ceiling, because an unbounded verification loop is the single most common way I’ve seen people burn real money on nothing. I ran one of these once without a cap during a translation task and watched the evaluator find increasingly pedantic objections for forty-one rounds before I killed it.
# verification_loop.py
# Evaluator-optimizer with a hard iteration cap, running against a local
# Ollama model. No API key, no per-token billing while you develop this.
#
# ollama pull llama3.1
# ollama serve
# pip install openai
from openai import OpenAI
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
MODEL = "llama3.1"
MAX_ROUNDS = 3
def generate(task: str, feedback: str | None = None) -> str:
prompt = task if not feedback else f"{task}\n\nAddress this feedback:\n{feedback}"
resp = client.chat.completions.create(model=MODEL, messages=[{"role": "user", "content": prompt}])
return resp.choices[0].message.content
def evaluate(task: str, candidate: str) -> tuple[bool, str]:
resp = client.chat.completions.create(
model=MODEL,
messages=[{
"role": "user",
"content": f"Task: {task}\nCandidate: {candidate}\n\nReply PASS or FAIL on line one, then one sentence why.",
}],
)
text = resp.choices[0].message.content
return text.strip().upper().startswith("PASS"), text
def verify_and_optimize(task: str) -> str:
candidate = generate(task)
for round_num in range(1, MAX_ROUNDS + 1):
passed, feedback = evaluate(task, candidate)
print(f"round {round_num}: {'PASS' if passed else 'FAIL'} -- {feedback[:80]}")
if passed:
return candidate
candidate = generate(task, feedback)
return candidate # ceiling hit, return the best attempt rather than loop forever
if __name__ == " __main__":
print(verify_and_optimize("Write one sentence explaining TCP handshakes for a non-technical reader."))
Point that same base_url at https://api.openai.com/v1 or https://api.anthropic.com/v1 with a real key and nothing else in the function bodies changes. The cap is the whole point of this layer existing separately from the taxonomy pattern it's built on: evaluator-optimizer describes the shape, the harness layer is what stops the shape from running forever.
Layer five: context pipelines
Context engineering is the layer that gets skipped most often because it doesn’t look like a bug until your token bill does. The core discipline is three verbs: select what’s relevant, compress what’s verbose, isolate what shouldn’t bleed between subtasks. Here’s a small, honest version of a context pipeline that does all three against a token budget instead of just appending everything to a list forever:
# context_pipeline.py
# Selects, compresses, and budgets context before it ever reaches the model.
def estimate_tokens(text: str) -> int:
return max(1, len(text) // 4) # rough heuristic, good enough for budgeting
def build_context(history: list[dict], project_notes: str, budget_tokens: int = 6000) -> list[dict]:
fixed = [{"role": "system", "content": project_notes}]
fixed_cost = estimate_tokens(project_notes)
# select: keep the most recent turns first, since they're usually most relevant
selected = []
running_cost = fixed_cost
for turn in reversed(history):
cost = estimate_tokens(turn["content"])
if running_cost + cost > budget_tokens:
# compress: summarize anything older that didn't fit, instead of dropping it silently
remaining = history[: history.index(turn) + 1]
summary = f"[{len(remaining)} earlier turns summarized: " + \
"; ".join(t["content"][:40] for t in remaining[-3:]) + "]"
selected.insert(0, {"role": "system", "content": summary})
break
selected.insert(0, turn)
running_cost += cost
return fixed + selected
if __name__ == " __main__":
fake_history = [{"role": "user", "content": f"turn {i}: " + "x" * 200} for i in range(50)]
ctx = build_context(fake_history, "Project uses FastAPI and Postgres.", budget_tokens=2000)
print(f"{len(ctx)} context entries, ~{sum(estimate_tokens(c['content']) for c in ctx)} tokens")
Every one of the five layers above is independent of which model you point it at, which is the whole point of building them yourself once instead of trusting that whichever harness product you picked implements all five well. Some don’t. That’s what the next section is about.
Three open-source harnesses that disagree with each other on purpose
Once I had the five layers built by hand, I went and installed three real, shipping, open-source harnesses that each made a different bet on where complexity should live. None of them is wrong. They’re three different answers to the same design question.
Everything is a plugin. DeepSeek Harness (dsh, MIT licensed) is built on Cordis, a plugin meta-framework where the agent loop, the tool scheduler, the sandbox, and even the SDK are each a plugin implementing a shared Service interface, resolved and hot-swappable through a dependency graph rather than a hand-written boot sequence. It's genuinely elegant engineering. It's also a 359-megabyte install with a default model catalog of exactly two models, and if you point it at a real repo that has both a CLAUDE.md and an AGENTS.md that have drifted apart, it loads both files into every single turn with no deduplication, no caching discount on the duplication. Read the plugin code before you trust what's actually being called; the getting-started docs undersell how much lives in there.
A tiny embeddable binary. Vercel Labs’ fx (Apache 2.0, on GitHub as vercel-labs/fx) is the opposite bet entirely: a coding agent harness written in Zig, compiled to roughly 7.8 MiB, with a hard CI-enforced ceiling on binary growth and a stated cold-start time in the microseconds for accepting input. It routes model calls through Vercel's AI Gateway, which charges zero markup on top of the underlying provider's list price, plus direct auth paths for OpenAI Codex and xAI Grok. It extends through skills, MCP servers, and subagents, three mechanisms instead of dsh's one universal plugin interface. What it doesn't have, as of when I checked, is a native Ollama provider entry the way Codex does; you can point it at any OpenAI-compatible endpoint through the gateway config, but it's not a first-class documented path the way it is for Codex.
Isolation as a first-class citizen. Alibaba’s OpenSandbox (Apache 2.0, alibaba/OpenSandbox) isn't a full agent harness in the sense of owning a reasoning loop, and I want to be upfront about that rather than force a comparison that doesn't fit. What it is: a sandbox platform purpose-built to sit underneath a harness, giving you a unified API across Docker, Kata Containers, gVisor, and Firecracker microVMs, with SDKs in Python, Java, Go, and a few others, plus MCP server support so any harness that speaks MCP can hand it code to execute. If your harness of choice treats sandboxing as an afterthought, this is the honest self-hosted way to bolt a real isolation boundary on afterward instead of trusting a bare docker run.
+------------------+------------------------+------------------------+------------------------+
| Dimension | DeepSeek Harness (dsh) | fx (Vercel Labs) | OpenSandbox (Alibaba) |
+------------------+------------------------+------------------------+------------------------+
| What it is | Full agent harness, | Full agent harness, | Sandbox/execution layer,|
| | plugin-native core | tiny native binary | not a reasoning loop |
+------------------+------------------------+------------------------+------------------------+
| Local models | No native provider, | No native Ollama entry, | N/A, runs whatever code |
| | wire a custom OpenAI- | route any OpenAI-compat | your harness sends it, |
| | compatible endpoint | endpoint via Gateway | model-agnostic by design|
+------------------+------------------------+------------------------+------------------------+
| Plugin system | Core design principle, | Skills + MCP + | MCP server support, |
| | Cordis DI, ctx keys, | subagents, three | otherwise a plain API, |
| | reversible effects | separate mechanisms | not plugin-first |
+------------------+------------------------+------------------------+------------------------+
| Context handling | Loads CLAUDE.md and | Permission + session | Not applicable, this is |
| | AGENTS.md in full, no | management, no public | infra a harness calls |
| | dedup if they diverge | detail on compaction | into, not context-aware |
+------------------+------------------------+------------------------+------------------------+
| Rough cost profile | Cheap DeepSeek models | Zero gateway markup, | Free and self-hosted, |
| | by default, but router | pay provider list price, | your cost is compute: |
| | overhead and duplicate | cost is whatever model | a Firecracker VM runs |
| | context add unseen tax | you choose to call | for pennies per hour |
+------------------+------------------------+------------------------+------------------------+
The honest takeaway from putting these three side by side isn’t “pick a winner.” It’s that “harness” bundles at least two genuinely separate concerns, the reasoning loop and the execution environment, and these three projects each chose to specialize in a different corner of that space rather than trying to own all of it.
Claude Code’s graph runtime as a case study
I want to close on this because it’s the most concrete, verifiable answer I found to the graph orchestration row in the taxonomy table, and because Anthropic actually documents the mechanism rather than leaving it as a marketing claim. Claude Code ships what its docs call dynamic workflows: a JavaScript script, written by Claude itself for the task you describe, that a separate runtime executes in the background, orchestrating up to sixteen concurrent subagents and up to a thousand total per run, while your main session’s context window never sees the intermediate results, only the final answer.
The framing that clicked for me, from a piece by Lakshman Sai arguing that most developers are sitting on this capability without opening it, is that the orchestration layer itself costs zero model tokens, because it’s code, not a conversation. Reading the actual mechanism confirms the claim rather than just repeating it: the script’s control flow, its loops, its branching, its pipeline() and parallel() calls, execute as plain JavaScript in an isolated runtime. The only thing that costs tokens is an agent() call, which spawns an actual subagent that talks to a model. That distinction is the first of the two cost levers, and it's a real architectural choice, not a rounding-error detail: free edges, paid nodes. You can add as much branching, retry logic, and conditional routing as the task needs without it costing anything beyond the compute to run a script, and your bill only moves when a node in that graph actually calls a model.
The second lever is per-node model tiering, and it’s spelled out plainly in Anthropic’s own documentation: every agent in a workflow uses your session’s default model unless the script explicitly routes a specific stage to a different one. In practice this means a workflow can run its planning or synthesis node on a stronger, more expensive model while every file-level worker node underneath it runs on something cheaper, and I’ve seen the same shape described independently in the broader graph-engineering discourse this year as an “advisor-orchestrator” pattern, reporting roughly 92 percent of top-tier quality at roughly 63 percent of the cost by keeping the expensive model at the coordination layer and pushing cheap models down to the repetitive execution layer. I haven’t reproduced that specific percentage myself against a controlled benchmark, so I’m flagging it as a plausible, directionally-consistent claim rather than a verified number, but the mechanism it depends on, per-stage model overrides, is real and documented, not speculative.
Before you reach for any of this, the actual question is whether your loop deserves to become a graph at all, and after building a handful of both, here’s the four-question test I now run before deciding, adapted from a decision framework I found genuinely useful in the “graph versus loop” discourse and confirmed against how Claude Code’s own workflow primitives are shaped:
+---+---------------------------------------------------------------+
| Q | Question |
+---+---------------------------------------------------------------+
| 1 | Does the work split into pieces that genuinely need separate |
| | context windows and tools, not one agent switching hats mid- |
| | transcript? |
+---+---------------------------------------------------------------+
| 2 | Is there real fan-out and fan-in, work that runs in parallel and |
| | gets merged back, not just a sequence you drew as boxes? |
+---+---------------------------------------------------------------+
| 3 | Do you want the control flow itself, the branching and the |
| | intermediate state, to live as inspectable code you can read, |
| | diff, and rerun, instead of buried in one agent's turn-by-turn |
| | judgment calls? |
+---+---------------------------------------------------------------+
| 4 | Will you run this same shape of task again, so that the cost of |
| | writing and saving the orchestration actually pays for itself? |
+---+---------------------------------------------------------------+
Zero or one “yes” and you’ve just renamed a loop. Two or three and you’re genuinely doing graph composition. All four and it’s worth saving as a reusable script rather than a one-off. Claude Code’s own progress view actually enforces a version of question four for you: it flags any run scheduling more than 25 agents or projecting past 1.5 million tokens with a visible warning, precisely because most tasks don’t clear that bar and shouldn’t be silently allowed to.
The three topologies I keep reaching for, matched to the actual runtime primitives Claude Code exposes, agent() for a single call, pipeline() for one-per-item sequencing, and parallel() for simultaneous fan-out, cover almost everything I've needed: a straight pipeline for a migration where each file transforms independently, a fan-out-then-synthesize shape for research or review where many sources get read in parallel and one node reconciles them, and a loop-until-done shape for the "keep fixing until the type check passes" case, capped at a fixed number of rounds for exactly the same reason the verification loop in layer four needed a ceiling. None of these three are exotic. What's new is that the runtime enforcing them ships inside a tool most people already have open, and the orchestration is free to write as long as you're honest with yourself about whether the underlying task actually needed a graph in the first place.
Where I landed
Seven names in a taxonomy table look interchangeable until you’ve built the five layers underneath them by hand and watched exactly where each one breaks. The layers don’t care which of the nine patterns you’re running, they’re the plumbing every pattern depends on regardless of its shape. The three open-source harnesses I installed each bet on a different corner of that plumbing being the one worth specializing in, and none of those bets is wrong, they’re just not the same bet. And the graph runtime sitting inside a tool a lot of us already use every day turned out to be less about a new pattern and more about making two cost levers, free edges and tiered models, into something you can actually see and control instead of something baked invisibly into a vendor’s black box.
If there’s one thing worth taking away past all the code: the word “harness” was never going to mean one thing, and pretending it does is how you end up comparing a system prompt to a 450,000-line codebase and calling it a fair fight. Ask which of the five layers a given product actually implements, ask which of the nine patterns it’s really running under the hood, and the overloaded word stops mattering nearly as much as the specific, checkable answers underneath it.
Tags: agent-harness, ai-agents, react-agents, claude-code, graph-orchestration, llm-engineering, open-source-ai
Top comments (0)