A source-code study of eleven coding agents gave me a vocabulary I didn’t know I was missing. So I turned it on my own projects.
The sentence that stopped me
Last week I was explaining an incident to a colleague. One of our agents had gotten stuck in a tool-calling loop overnight and burned through a chunk of budget before anyone noticed. Halfway through my explanation I said “the agent framework retried the call” and then paused, because that wasn’t true.
There was no agent framework in that code path. It was a raw SDK call with a while loop I had written myself, eight months earlier, on a Friday afternoon. The retry was mine. The missing iteration cap was mine. The "framework" I was blaming was me.
That small moment bothered me more than the incident did. I realized I had been using “the agent framework” as a catch-all for at least three very different things in my stack:
- Claude Code, which I use every day as a coding agent
- Microsoft Agent Framework, which runs a couple of our .NET agents
- A handful of direct Anthropic SDK integrations, some using the tool runner helper, some using a loop I wrote by hand
Those three things fail in completely different ways. They put the control loop in different hands. They store state in different places. And I had been talking about them as if they were interchangeable.
The same week, I read Jaroslaw Wasowski’s piece on Level Up Coding about choosing a coding agent’s runtime. It is built around a new arXiv paper (2609.00006) by Paul Barbaste and co-authors, a source-code anatomy of eleven production coding harnesses: Claude Code, Codex CLI, Gemini CLI, Mistral Vibe, OpenHands, Aider, Mini-SWE-Agent, Hermes, Pi, OpenCode, and OpenClaw, with Databricks’ Omnigent analyzed as a contrast point. Wasowski’s headline reading of it stuck with me: across roughly four million lines of Python, TypeScript, and Rust, none of those agents imports a general-purpose agent framework at runtime. They all own their loop.
That was the push I needed. If the people building the most heavily used coding agents in the world all chose to own the loop, I wanted to know why I had made the choices I made. So I did a small audit of my own stack using the same lens.
Three words that are not synonyms
The paper’s core definition is simple. An agent is a model plus a harness, and the harness is everything except the model: the loop, the tools, context management, safety controls, orchestration, and extension surfaces. It maps seven subsystems that every harness has to take a position on, even if that position is “we don’t do this.”
Wasowski’s framing on top of that is what made it practical for me. The question is not “which framework should I pick.” The question is who owns the loop, and where does state have to survive a crash. From that, three categories fall out naturally:
Harness. An opinionated, batteries-included runtime. You bring the model, the instructions, and your domain tools. The harness brings the loop, planning, context compaction, approvals, memory, and telemetry. You configure it, but you don’t write the loop. Claude Code is the obvious example. You shape its behavior through CLAUDE.md, skills, hooks, and permissions, but you are a guest in its loop.
Framework. You own the loop, or at least the shape of it. The framework gives you primitives: graphs, nodes, checkpoints, state reducers, workflow steps. LangGraph lives here. So does the workflow side of Microsoft Agent Framework. You decide what happens after each step. The framework decides how steps are wired, persisted, and resumed.
SDK. You own everything. The SDK gives you a typed API call and maybe a small helper. Anthropic’s tool runner is about as far as it goes: it will run the tool loop for you, but it has no opinion about memory, approvals, persistence, or what happens when your process dies halfway through.
Here is the version I ended up pinning above my desk:
+-------------+----------------------+-------------------------+--------------------------+
| | HARNESS | FRAMEWORK | SDK |
+-------------+----------------------+-------------------------+--------------------------+
| Loop owner | The runtime | You, via its primitives | You, entirely |
| State | Runtime decides | You declare, it persists| You build it (or not) |
| Safety | Built in, configured | Pluggable | Your code or nothing |
| Typical | Claude Code, | LangGraph, MAF | Anthropic SDK, |
| example | MAF AsHarnessAgent | workflows | tool runner, raw loop |
| Fails like | Opaque, "why did it | Graph/state bugs, | Missing guardrails you |
| | do that?" | version churn | forgot to write |
| Best for | Open-ended tasks, | Long-running, resumable | Narrow, bounded calls |
| | humans in the loop | pipelines | inside a bigger app |
+-------------+----------------------+-------------------------+--------------------------+
One wrinkle worth calling out: these are not product categories, they are roles. Microsoft Agent Framework plays two of them. Its workflow APIs are a framework in the classic sense. But it now also ships an agent harness, exposed through AsHarnessAgent, which is exactly the batteries-included runtime described above. The same NuGet dependency can put you in either column depending on which API you call. That confused me at first, and I think it confuses a lot of .NET teams.
My three integration points
I picked three real places in my stack where an LLM does something more than one-shot text generation, and asked the same two questions of each: who owns the loop, and what state must survive a crash.
1. The customer-facing assistant: I picked a harness
This is a support assistant that answers questions about device connectivity, reads account data, and can open a ticket or trigger a diagnostic action on the customer’s behalf.
When I listed what it needed, the list was long: tool calling, a plan for multi-step questions, conversation history that survives restarts, context compaction for long chats, file access scoped to one folder of generated reports, human approval before anything with side effects, and traces I can show someone after a complaint.
I could build all of that. I have built parts of it before, badly. But none of it is what makes this assistant valuable. What makes it valuable is the domain tools and the boundaries around them.
So this one runs on the Microsoft Agent Framework harness. Bruno Capuano’s MafClaw series on the .NET blog is a good walkthrough of the pattern, and the starting point really is this small:
using Azure.AI.Projects;
using Azure.Identity;
using Microsoft.Agents.AI;
using Microsoft.Extensions.AI;
var endpoint = Environment.GetEnvironmentVariable("FOUNDRY_PROJECT_ENDPOINT")!;
var model = Environment.GetEnvironmentVariable("FOUNDRY_MODEL") ?? "gpt-4.1-mini";
IChatClient chatClient =
new AIProjectClient(new Uri(endpoint), new AzureCliCredential())
.GetProjectOpenAIClient()
.GetResponsesClient()
.AsIChatClient(model);
var workingDirectory = Path.Combine(AppContext.BaseDirectory, "working");
Directory.CreateDirectory(workingDirectory);
AIAgent agent = chatClient.AsHarnessAgent(new HarnessAgentOptions
{
// File tools are confined to this folder. The model never sees the rest of the disk.
FileAccessStore = new FileSystemAgentFileStore(workingDirectory),
// Reads go through silently. Writes and side effects still need a human.
ToolApprovalAgentOptions = new ToolApprovalAgentOptions
{
AutoApprovalRules = [FileAccessProvider.ReadOnlyToolsAutoApprovalRule]
},
ChatOptions = new ChatOptions
{
Instructions = """
You are a support assistant for device connectivity questions.
Use get_device_status before answering anything about a specific device.
Never promise a fix. Offer to open a ticket instead.
""",
Tools = [DeviceTools.GetDeviceStatus, DeviceTools.RequestDiagnosticRun]
}
});
Console.WriteLine(await agent.RunAsync("Why is device 4417 offline?"));
And the tools are ordinary C#. The one with side effects gets wrapped so the model can request it but never run it directly:
using System.ComponentModel;
using Microsoft.Extensions.AI;
public static class DeviceTools
{
[Description("Gets the current connectivity status for a device.")]
public static string GetDeviceStatusById(
[Description("Device identifier, e.g. 4417")] string deviceId)
=> deviceId switch
{
"4417" => "4417: offline since 03:12 UTC, last seen on cell 2231 (mock)",
_ => $"{deviceId}: online (mock)"
};
[Description("Requests a remote diagnostic run on a device.")]
public static string RequestDiagnosticRunById(string deviceId)
=> $"Diagnostic run queued for {deviceId} (mock)";
public static AIFunction GetDeviceStatus { get; } =
AIFunctionFactory.Create(GetDeviceStatusById, "get_device_status");
public static AIFunction RequestDiagnosticRun { get; } =
new ApprovalRequiredAIFunction(
AIFunctionFactory.Create(RequestDiagnosticRunById, "request_diagnostic_run"));
}
A quick note on packages: the harness APIs were still prerelease when I wrote this, so the exact package set moves. I install with dotnet add package Microsoft.Agents.AI --prerelease plus the Azure AI Projects packages, and I check the MafClaw sample repo (aka.ms/mafclaw/repo) for the current versions before I trust anything I remember.
If you don’t want a cloud model for this at all , the harness does not care where the IChatClient comes from. For local development I swap the first block for Ollama through OllamaSharp, which implements IChatClient directly:
# Run Ollama locally (or: brew install ollama && ollama serve)
docker run -d --name ollama -p 11434:11434 -v ollama:/root/.ollama ollama/ollama
docker exec ollama ollama pull qwen3:8b
dotnet add package OllamaSharp
using OllamaSharp;
IChatClient chatClient = new OllamaApiClient(
new Uri("http://localhost:11434"), "qwen3:8b");
// Everything below, AsHarnessAgent(...) included, stays exactly the same.
Pick a local model that actually supports tool calling. Small models will happily ignore your tools and answer from vibes, and the harness can’t fix that for you.
The part that convinced me the harness was the right tier was the approval flow. In one of the MafClaw sessions someone asked what happens if the user never answers an approval prompt. The answer became a whole sample with timeouts and bounded retries. Silence is not consent. That is the kind of edge case I would not have thought about until production taught me, and a harness that has already been taught it is worth a lot.
What I gave up: transparency. When the assistant does something odd, the first question is always “was that the model, or the harness?” I now keep OpenTelemetry tracing on from day one, not as an afterthought, because a harness without traces is a black box you are responsible for.
2. The internal batch agent: I picked a framework
The second integration is not a conversation. It is a nightly job that reads a few thousand support tickets, classifies them, pulls related device telemetry, drafts a summary per customer segment, and writes a report that product managers read over coffee.
Nobody is chatting with it. Nobody approves individual steps. What matters is completely different:
- If it crashes at ticket 2,300, it must resume from ticket 2,301, not from zero
- Each stage must be inspectable on its own when a summary looks wrong
- Token spend per run needs a hard ceiling
This is the “where must state survive a crash” question, and here the answer is unambiguous: at every step boundary.
A harness is the wrong shape for this. Harnesses are built around an open-ended conversational loop where the model decides what comes next. My batch job has a fixed shape. Classify, enrich, summarize, write. The model should not decide the order of stages. I should.
So this runs on Microsoft Agent Framework workflows, the framework side of the same library. Each stage is an executor, the edges are declared, and checkpoints persist state between them. The model does real work inside each node, but the graph is mine.
I considered LangGraph seriously. It is arguably more mature for exactly this pattern. What tipped it was boring: the rest of this pipeline is .NET, the team reads C#, and the telemetry already flows through the same OpenTelemetry setup as everything else. Framework choice is rarely about the framework.
What I gave up: simplicity, and some stability. Framework APIs in this space move fast. I have rewritten graph wiring code twice this year because of breaking changes, which is the tax you pay for renting someone else’s primitives.
3. The narrow helper: I picked an SDK (the second time)
The third integration is tiny. Inside a larger ASP.NET Core application there is a form where an operator types a free-text description of a problem, and a helper turns that into a structured filter: device group, time window, error class. It has exactly one tool, which looks up whether a device group name is valid. It runs maybe a few hundred times a day.
The loop here is at most two or three turns. There is no state worth persisting. If it fails, the operator types the filter by hand, which is what they did before.
This is SDK territory. Here is roughly what it looks like, written against IChatClient so it runs the same against Anthropic in production and Ollama on my laptop:
using System.ComponentModel;
using Microsoft.Extensions.AI;
public sealed class FilterHelper(IChatClient client)
{
private const int MaxTurns = 3;
[Description("Returns true if the device group name exists.")]
private static bool IsValidGroup(string groupName) =>
groupName is "eu-fleet" or "us-fleet" or "test-bench"; // real lookup goes here
private readonly AIFunction _groupTool =
AIFunctionFactory.Create(IsValidGroup, "is_valid_group");
public async Task<string?> BuildFilterAsync(string input, CancellationToken ct = default)
{
List<ChatMessage> messages =
[
new(ChatRole.System,
"Convert the request into JSON with keys group, from, to, errorClass. " +
"Validate the group with is_valid_group. Reply with JSON only."),
new(ChatRole.User, input)
];
var options = new ChatOptions { Tools = [_groupTool], MaxOutputTokens = 300 };
for (var turn = 0; turn < MaxTurns; turn++)
{
ChatResponse response = await client.GetResponseAsync(messages, options, ct);
messages.AddMessages(response);
var calls = response.Messages
.SelectMany(m => m.Contents)
.OfType<FunctionCallContent>()
.ToList();
if (calls.Count == 0)
return response.Text; // done, model answered
foreach (var call in calls)
{
var result = await _groupTool.InvokeAsync(
new AIFunctionArguments(call.Arguments), ct);
messages.Add(new ChatMessage(ChatRole.Tool,
[new FunctionResultContent(call.CallId, result)]));
}
}
return null; // hit the cap, caller falls back to the manual form
}
}
Every line of that loop is mine, including the MaxTurns cap that the Friday-afternoon version of me forgot. That is the whole point. There is nothing here I can't see.
If you are in Python, Anthropic’s tool runner gets you the same shape with even less code:
from anthropic import Anthropic, beta_tool
client = Anthropic() # reads ANTHROPIC_API_KEY
@beta_tool
def is_valid_group(group_name: str) -> bool:
"""Return True if the device group name exists."""
return group_name in {"eu-fleet", "us-fleet", "test-bench"}
runner = client.beta.messages.tool_runner(
model="claude-sonnet-5",
max_tokens=300,
tools=[is_valid_group],
messages=[{"role": "user", "content": "errors in eu-fleet since Monday"}],
)
final = runner.until_done()
print(final.content[0].text)
Just remember that the runner runs the loop, but it does not bound it for you in any way that fits your budget. Put your own limit around it.
The wrong choice, and what it cost me
I said “the second time” above because the first version of that helper was not built on an SDK. It was built on a framework.
At the time we were adopting a graph-based framework for the batch pipeline, and I was excited about it. So when the filter helper came up, I reached for the same tool. I modeled a three-turn interaction as a graph: a parse node, a validate node, a conditional edge back to parse, a checkpointer (in-memory, because what was I going to persist?), and a state schema with five fields for what was really a list of messages.
It worked. It also cost me in ways that took months to add up:
- Latency. The framework’s per-step overhead was small, but the helper sits on a form submit. Users noticed a sluggish feel that the raw call didn’t have once I measured them side by side.
- Tokens. My graph’s state carried the full intermediate structure into every node’s prompt. The raw version sends only the messages it needs. The prompt got noticeably smaller once I rewrote it.
- Upgrades. Two framework version bumps broke that helper even though it used almost none of the framework’s features. Each time, someone had to relearn the graph API to fix a component that should have been forty lines of plain C#.
- Debugging. When the helper returned garbage, the stack trace went through the framework’s scheduler. For a two-turn loop, that is absurd.
The rewrite took one afternoon. The code shrank to what you saw above. Nothing got worse.
The lesson I took from it is not “frameworks are bad.” The batch agent is much better for being on one. The lesson is that I chose a tier based on what I was excited about, not on who needed to own the loop. A framework earns its keep when state has to survive something. When there is no state worth surviving, you are paying rent for an empty apartment.
Elif Sude Gökay makes a point in her piece on harness engineering that fits here: AI can make code cheap, but it does not automatically make trustworthy change cheap. The runtime choice is part of what makes change trustworthy, or not. Picking the heaviest tier by default is a way of making every future change more expensive.
A fair warning about the audit itself
I want to be careful not to turn one paper into scripture, and I think the paper would agree.
First, it is a source-code study, not a benchmark. It tells you how these harnesses are built. It does not tell you which one produces better outcomes on your tasks, and nothing in it measures that directly.
Second, the snapshots are snapshots. The authors pinned specific versions and even compared eight of them across one quarter, which is genuinely useful because it shows how fast this space converges and copies itself. But it also means some of what they describe may already be different in the version you are running today. Claude Code in particular is closed-source, and as far as I can tell the analysis of it is much harder for an outsider to reproduce than, say, the Aider or Mini-SWE-Agent sections. I read those parts with more caution.
Third, the “zero frameworks at runtime” finding is about coding harnesses built by teams whose whole product is the harness. Of course they own the loop. That is their job. It does not follow that your internal ticket pipeline should hand-roll its own checkpointing. The finding tells you what vendors do, not what you should do.
So I treat the taxonomy the way I treat a good architecture pattern name: as a thinking tool that makes conversations shorter and decisions more honest. Not as a ranking.
The checklist I use now
After the audit, I wrote down the questions I actually ask before the next integration. They are short on purpose.
1. Who should decide what happens next: the model, or me? If the model should drive (open-ended, conversational, unknown number of steps), lean harness. If I should drive (fixed stages, known shape), lean framework or SDK.
2. What state must survive a crash, and where? If the answer is “nothing,” you almost certainly want an SDK. If the answer is “every step boundary,” you want a framework with real checkpointing. If the answer is “the conversation and some user memory,” a harness usually has that covered.
3. Is a human in the loop for side effects? Approval flows are surprisingly hard to get right (timeouts, retries, sticky denials). If you need them and they aren’t your product, a harness that ships them is worth its opacity.
4. How many turns, realistically? One to three turns with one or two tools is SDK territory. Write the loop, put a cap on it, move on.
5. Can I trace it on day one? Whatever tier you pick, if you can’t see every model call and tool call, you will end up blaming the wrong layer. Ask me how I know.
6. Can it run locally? If the runtime only works against one hosted model, your dev loop and your tests get expensive. Anything built on IChatClient (or an OpenAI-compatible endpoint in Python) can point at Ollama for development. I now treat that as a requirement, not a nice-to-have.
7. Am I picking this because it fits, or because I’m already using it elsewhere? This is the one that would have saved me the most. Reuse is a good reason to pick between two options that both fit. It is a bad reason to pick an option that doesn’t.
Where I landed
My stack didn’t change much after this exercise. The chatbot stayed on a harness, the batch job stayed on a framework, and the helper had already been fixed. What changed is the vocabulary. When something breaks now, I can say which layer owns the loop, and that usually tells me where to look.
And when someone on the team says “the agent framework did it,” I ask the question that embarrassed me a few weeks ago: which one?
Sources
- Jaroslaw Wasowski, “Harness vs. Framework vs. SDK: How to Choose a Coding Agent’s Runtime,” Level Up Coding: https://levelup.gitconnected.com/harness-vs-framework-vs-sdk-how-to-choose-a-coding-agents-runtime-2ac776994c3a
- Paul Barbaste et al., “Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents: A Source-Code Study of Eleven Systems,” arXiv 2609.00006: https://arxiv.org/abs/2609.00006
- Bruno Capuano, “Build Your Own AI Agent Harness in C#, the MafClaw Live Series,” .NET Blog: https://devblogs.microsoft.com/dotnet/build-your-own-ai-agent-harness-in-csharp-the-maf-claw-live-series/
- Elif Sude Gökay, “Harness Engineering: When AI Writes the Code, What Are Engineers Actually Engineering?”: https://ai.plainenglish.io/harness-engineering-when-ai-writes-the-code-what-are-engineers-actually-engineering-4d45e21c1777
Tags: AI Agents, Harness Engineering, Dotnet, CSharp, Software Architecture, Claude, LLM, Microsoft Agent Framework, Ollama
Strongest 5 for Medium: AI Agents, Harness Engineering, Dotnet, Software Architecture, LLM
Top comments (0)