I built the same memory system twice, once for a project with three people and once for a codebase eleven teams touch, and I’m no longer sure both projects should use the same answer
I want to open with a disagreement I have not resolved, because I think pretending it’s settled is the dishonest version of this article.
A few weeks ago I watched Theo, who runs t3.gg and spends a lot of his time in the guts of coding agents, walk through an audit of his own Claude Code memory setup on stream. He found 45 memory files sitting in his main project. Twenty six of them had never been read back by the agent, not once. Memories were being written roughly three times as often as they were ever retrieved. His conclusion wasn’t “tune the retrieval.” It was “turn memory off,” and he did, across his whole fleet of agents, and told everyone watching to do the same. His argument, in short: code is the ground truth, a good grep over the actual repository beats a curated memory store that's mostly write-only, and for most teams the memory system is theater that costs tokens and adds a place for stale nonsense to hide.
I read that and felt seen, because I have a folder full of memory files in a side project that I would bet money the agent has never opened either.
Then I read the other side of it, and it’s not a strawman. Kapil Viren Ahuja, writing about this exact tension, makes a claim I can’t wave off: he agrees the minimalists are right, for individuals and small teams, and then argues that if you run engineering at an enterprise and follow that advice as written, your project will not survive the first ten people who join it. His framing, which stuck with me, is that in a brownfield enterprise codebase “what they delete is what saves you.” Not because curated memory is inherently superior to grep, but because the thing grep is searching, the actual code, is not a reliable narrator once dozens of engineers across several years have touched it.
Both of these people are describing something true. I spent the better part of a month trying to figure out where the line between them actually sits, by building the same kind of thing for two very different codebases, and this article is the honest version of what I found, including the part where I had to admit my test was flaky and fix it.
Camp one: memory is mostly theater, and grep already won
The minimalist case is stronger than “vector databases are overhyped,” which is the version of this argument I expected to find and didn’t. It’s a claim about what a language model is actually good at.
A model has seen an enormous amount of pretraining data involving grep, find, reading a file, walking a directory tree. It has seen comparatively little data involving the specific shape of your memory API, your bespoke fact schema, or the particular way your curated context gets assembled before a call. When an agent reaches for a tool, it's working from familiarity with that tool's shape, not from some abstract judgment about which retrieval method is theoretically best. Grep isn't a good retrieval algorithm in the information-retrieval sense. It's an extremely familiar one, and familiarity is what lets an agent use it correctly on the first try instead of half-remembering a query syntax from a different system.
Theo’s number is the concrete version of that argument. If you build a memory store and 58 percent of what goes into it (26 of 45 files) never gets read back out, you haven’t built a memory system, you’ve built a write-only log with extra steps and extra token cost on every write. The honest read of that audit isn’t “memory is bad,” it’s “this particular memory system had no retrieval discipline, so it decayed into noise, and the code underneath it was current enough that grep found the real answer faster anyway.”
For a small team on a codebase one or two people can hold in their heads, I think that’s just correct. The code is current. The conventions are whatever the last commit says they are, because the same small group of people wrote all of it and would notice if a comment lied. There’s no tribal knowledge locked in a departed engineer’s head, because there’s no departed engineer yet. In that world, a curated memory layer is solving a problem you don’t have, and every fact you cache is a fact that can go stale the moment someone refactors and forgets to update your memory store alongside the code.
Camp two: the code stops being a reliable narrator at scale
Here’s where I stopped being able to side entirely with Theo, and it happened on a codebase that wasn’t mine.
A team I was helping had a service where the README said "we deploy to staging via the pipeline in deploy/legacy," a comment three folders over said // TODO: migrate off legacy deploy, use deploy/v2, and deploy/v2 had been the actual production path for over a year. Nobody had deleted the old folder because deleting it required a conversation with a team that had since been reorganized twice. Grep over that repository doesn't return one truth. It returns three artifacts that each look authoritative and disagree with each other, and an agent has no principled way to know which one reflects what actually happens when code ships, because syntactically all three are just as valid as the others.
This is the shape of Ahuja’s argument, and it’s the one I now take seriously: a brownfield enterprise codebase accumulates dead code paths, comments that describe an intention nobody followed through on, and conventions that drifted across however many engineers touched the repository over however many years, and none of that shows up to grep as noise. It shows up as more code, indistinguishable in format from the code that’s actually true.
I ran my own small before-and-after on this, because I didn’t want to just take the claim on faith. I gave an agent a task, “scaffold the infrastructure-as-code for a new internal service,” on a real org’s repository, with nothing in its context beyond the codebase itself. The org is an Azure shop, has been for years, every existing service in that monorepo is provisioned through Bicep against Azure resources. The agent’s first attempt came back with Terraform targeting AWS. Not because anything in the repo said AWS, but because nothing in the repo said “we are an Azure shop and this is not a decision you get to make,” and in the absence of that constraint the model fell back on the statistically dominant pattern in its own training data, which skews AWS for generic infrastructure-as-code requests. I then wrote exactly one fact into a memory layer, “this organization provisions all cloud infrastructure on Azure via Bicep, this is a standard, not a preference,” loaded it ahead of the same prompt, and the agent’s next attempt came back correctly targeting Azure with Bicep. Same model, same repository, same prompt. The only variable that changed was one sentence of organizational memory that no file in the repository stated outright, because it didn’t need to be stated for a human who already knew it.
That’s the actual crux for me. Grep is unbeatable at finding what the code says. It has no mechanism at all for surfacing what the code assumes everyone already knows.
What actually belongs in memory, versus what belongs in the prompt
Before I get to the .NET implementation, I want to name a distinction that both camps actually agree on once you dig past the headline disagreement, because it’s what the practical section below is built on. It’s a framing I first saw laid out cleanly in a piece by Marwa Samy on what an AI agent should remember about a user, and it reframes the whole debate usefully: the question isn’t “how do I make the model remember everything,” it’s “what does my application actually need to remember.”
Conversation history and state are not the same thing, and conflating them is most of what makes memory systems bloat into Theo’s 45-file mess. Conversation history is the raw back-and-forth, every message, every hedge, every “actually, let’s change that.” State is the distilled, structured answer to “what does the agent currently know that it needs in order to keep helping this specific user or task.” If a user says, across three separate messages, “we’re on Azure,” “actually let’s keep this in West Europe for data residency,” and “we’ll probably want blue-green deployments eventually,” the conversation history is three messages of scattered context. The state is three fields: CloudProvider: Azure, Region: westeurope, DeploymentStrategy: blue-green (planned). Everything downstream, retrieval, prompting, memory hygiene, gets dramatically simpler once you stop trying to remember the conversation and start trying to remember the state.
That’s the design the .NET code below implements: a context provider that keeps raw history and extracted state in two separate places, extracts new facts after every turn, and merges them into existing state without duplicating or contradicting what’s already there.
Building it in .NET: a minimal AIContextProvider
The current Microsoft Agent Framework (the Microsoft.Agents.AI namespace) gives you an actual extension point for exactly this, AIContextProvider, an abstract class with two lifecycle hooks: one that runs before the model is invoked, where you inject retrieved context, and one that runs after, where you extract and persist whatever the turn taught you. Here's the shape, trimmed to what a minimal implementation actually needs.
using Microsoft.Agents.AI;
using Microsoft.Extensions.AI;
using System.Text.Json;
// The structured fact set we care about. Deliberately narrow: this is not
// a transcript, it's the handful of fields the agent actually needs on
// every future turn for this user or task.
public sealed class OrgMemoryState
{
public string? CloudProvider { get; set; }
public string? Region { get; set; }
public string? DeploymentStrategy { get; set; }
public Dictionary<string, string> AdditionalFacts { get; set; } = new();
}
public sealed class OrgMemoryContextProvider : AIContextProvider
{
private readonly IMemoryStore _store;
private readonly IChatClient _extractionClient;
private readonly string _userId;
public OrgMemoryContextProvider(IMemoryStore store, IChatClient extractionClient, string userId)
: base(null, null, null)
{
_store = store;
_extractionClient = extractionClient;
_userId = userId;
}
// Runs before the model is called. Loads current state, hands it to
// the model as a plain instruction, nothing exotic.
protected override async ValueTask<AIContext> ProvideAIContextAsync(
InvokingContext context, CancellationToken cancellationToken = default)
{
var state = await _store.LoadStateAsync(_userId, cancellationToken);
if (state is null)
return new AIContext();
var summary = $"""
Known organizational and user facts (treat as authoritative, do not contradict without asking):
- Cloud provider: {state.CloudProvider ?? "unknown"}
- Region: {state.Region ?? "unknown"}
- Deployment strategy: {state.DeploymentStrategy ?? "unknown"}
{string.Join('\n', state.AdditionalFacts.Select(kv => $"- {kv.Key}: {kv.Value}"))}
""";
return new AIContext
{
Messages = [new ChatMessage(ChatRole.System, summary)]
};
}
// Runs after the model responds. Extracts structured facts from this
// turn and merges them into whatever state already exists.
protected override async ValueTask StoreAIContextAsync(
InvokedContext context, CancellationToken cancellationToken = default)
{
if (context.InvokeException is not null)
return; // never learn from a failed turn
var turnText = string.Join(
"\n",
context.RequestMessages.Concat(context.ResponseMessages ?? [])
.Select(m => $"{m.Role}: {m.Text}"));
var extracted = await ExtractFactsAsync(turnText, cancellationToken);
if (extracted is null)
return; // nothing worth remembering from this turn
var existing = await _store.LoadStateAsync(_userId, cancellationToken) ?? new OrgMemoryState();
var merged = MergeState(existing, extracted);
await _store.SaveStateAsync(_userId, merged, cancellationToken);
}
// Extraction: turn free-form conversation into the structured shape
// above. This is the one call in the whole pipeline that talks to a
// model, and its only job is "did this turn state a fact I care about."
private async Task<OrgMemoryState?> ExtractFactsAsync(string turnText, CancellationToken ct)
{
var prompt = $"""
Read the exchange below. If it states or changes any of these facts,
return them as compact JSON with only the fields that were actually
stated: cloudProvider, region, deploymentStrategy, additionalFacts
(an object of free-form key/value pairs). If nothing relevant was
stated, return an empty JSON object: {{}}
Exchange:
{turnText}
""";
var response = await _extractionClient.GetResponseAsync(prompt, cancellationToken: ct);
var json = response.Text.Trim();
if (json is "{}" or "")
return null;
try
{
var doc = JsonDocument.Parse(json);
var state = new OrgMemoryState();
if (doc.RootElement.TryGetProperty("cloudProvider", out var cp))
state.CloudProvider = cp.GetString();
if (doc.RootElement.TryGetProperty("region", out var rg))
state.Region = rg.GetString();
if (doc.RootElement.TryGetProperty("deploymentStrategy", out var ds))
state.DeploymentStrategy = ds.GetString();
if (doc.RootElement.TryGetProperty("additionalFacts", out var af))
{
foreach (var prop in af.EnumerateObject())
state.AdditionalFacts[prop.Name] = prop.Value.GetString() ?? "";
}
return state;
}
catch (JsonException)
{
// A malformed extraction is a reason to skip this turn's write,
// not a reason to crash the conversation the user is in.
return null;
}
}
// Merge: a new fact for a key that already has a value overwrites it
// (the newer statement wins), a key that's new gets added, a key that
// isn't mentioned this turn is left completely alone.
private static OrgMemoryState MergeState(OrgMemoryState existing, OrgMemoryState incoming)
{
return new OrgMemoryState
{
CloudProvider = incoming.CloudProvider ?? existing.CloudProvider,
Region = incoming.Region ?? existing.Region,
DeploymentStrategy = incoming.DeploymentStrategy ?? existing.DeploymentStrategy,
AdditionalFacts = new Dictionary<string, string>(existing.AdditionalFacts)
.Concat(incoming.AdditionalFacts)
.GroupBy(kv => kv.Key)
.ToDictionary(g => g.Key, g => g.Last().Value),
};
}
}
public interface IMemoryStore
{
Task<OrgMemoryState?> LoadStateAsync(string userId, CancellationToken ct);
Task SaveStateAsync(string userId, OrgMemoryState state, CancellationToken ct);
}
Registering it is the part that made this click for me the first time I wired it up, because it’s just another item in the same list that holds tools and instructions:
AIAgent agent = chatClient.AsAIAgent(new ChatClientAgentOptions
{
Instructions = "You are an infrastructure assistant for this organization.",
AIContextProviders = [
new OrgMemoryContextProvider(memoryStore, extractionClient, userId: currentUserId)
],
});
The thing I got wrong the first time: I put the extraction call and the summarization-for-prompt call on the same IChatClient instance sharing conversation history, and ended up with the extraction step quietly picking up context bleed from the injected memory summary, so a fact would get "re-extracted" and re-written every single turn even when nothing new was said. The fix was boring, give the extraction call a clean, single-purpose prompt with no shared history, which is what the code above does.
Backing it with real storage: Azure Cosmos DB, and a local alternative
The part of this that actually matters for correctness, more than the storage technology you pick, is keeping two identifiers cleanly separate: the conversation or session ID, which is short-lived and dies with the conversation, and the user or tenant ID, which is long-lived and is what your memory is actually keyed on. Collapsing these two into one ID is the single most common bug I’ve seen in memory implementations, because it means every new conversation starts with total amnesia, which defeats the entire point.
+----------------------+------------------------+---------------------------+
| Identifier | Lifespan | What it answers |
+----------------------+------------------------+---------------------------+
| session_id / thread_id | One conversation, | "Which conversation is |
| | minutes to hours | this?" |
+----------------------+------------------------+---------------------------+
| user_id / tenant_id | Indefinite, survives | "Whose durable memory |
| | every future session | is this?" |
+----------------------+------------------------+---------------------------+
Azure Cosmos DB
Two containers, split by exactly that distinction. sessions holds the raw turn history, partitioned by sessionId, with a TTL so it expires on its own. userMemory holds the extracted state, partitioned by userId, with no TTL, because it's meant to outlive any single conversation.
using Microsoft.Azure.Cosmos;
public sealed class CosmosMemoryStore : IMemoryStore
{
private readonly Container _userMemory;
public CosmosMemoryStore(CosmosClient client, string databaseId)
{
_userMemory = client.GetContainer(databaseId, "userMemory");
}
public async Task<OrgMemoryState?> LoadStateAsync(string userId, CancellationToken ct)
{
try
{
var response = await _userMemory.ReadItemAsync<UserMemoryDocument>(
id: userId, partitionKey: new PartitionKey(userId), cancellationToken: ct);
return response.Resource.State;
}
catch (CosmosException ex) when (ex.StatusCode == System.Net.HttpStatusCode.NotFound)
{
return null;
}
}
public async Task SaveStateAsync(string userId, OrgMemoryState state, CancellationToken ct)
{
var doc = new UserMemoryDocument { Id = userId, UserId = userId, State = state };
await _userMemory.UpsertItemAsync(doc, new PartitionKey(userId), cancellationToken: ct);
}
private sealed class UserMemoryDocument
{
public string Id { get; set; } = default!;
public string UserId { get; set; } = default!;
public OrgMemoryState State { get; set; } = default!;
}
}
Container setup, once, via the SDK or the portal: partition key /userId on userMemory, no default TTL; partition key /sessionId on sessions, with DefaultTimeToLive set to something like 30 days so raw turns clean themselves up without a background job.
A self-hosted alternative: SQLite
If you don’t have an Azure subscription handy, or just want to run this on a laptop first, the identical two-table split works fine in SQLite with zero external services:
using Microsoft.Data.Sqlite;
using System.Text.Json;
public sealed class SqliteMemoryStore : IMemoryStore, IDisposable
{
private readonly SqliteConnection _conn;
public SqliteMemoryStore(string dbPath = "memory.db")
{
_conn = new SqliteConnection($"Data Source={dbPath}");
_conn.Open();
using var cmd = _conn.CreateCommand();
cmd.CommandText = """
CREATE TABLE IF NOT EXISTS user_memory (
user_id TEXT PRIMARY KEY,
state_json TEXT NOT NULL,
updated_at TEXT NOT NULL
);
CREATE TABLE IF NOT EXISTS session_turns (
session_id TEXT NOT NULL,
user_id TEXT NOT NULL,
turn_json TEXT NOT NULL,
created_at TEXT NOT NULL
);
""";
cmd.ExecuteNonQuery();
}
public async Task<OrgMemoryState?> LoadStateAsync(string userId, CancellationToken ct)
{
using var cmd = _conn.CreateCommand();
cmd.CommandText = "SELECT state_json FROM user_memory WHERE user_id = $uid";
cmd.Parameters.AddWithValue("$uid", userId);
var result = await cmd.ExecuteScalarAsync(ct);
return result is string json ? JsonSerializer.Deserialize<OrgMemoryState>(json) : null;
}
public async Task SaveStateAsync(string userId, OrgMemoryState state, CancellationToken ct)
{
using var cmd = _conn.CreateCommand();
cmd.CommandText = """
INSERT INTO user_memory (user_id, state_json, updated_at)
VALUES ($uid, $state, $now)
ON CONFLICT(user_id) DO UPDATE SET state_json = $state, updated_at = $now
""";
cmd.Parameters.AddWithValue("$uid", userId);
cmd.Parameters.AddWithValue("$state", JsonSerializer.Serialize(state));
cmd.Parameters.AddWithValue("$now", DateTimeOffset.UtcNow.ToString("O"));
await cmd.ExecuteNonQueryAsync(ct);
}
public void Dispose() => _conn.Dispose();
}
No container to spin up, no account to create, and it’s a drop-in IMemoryStore so you can run the exact same OrgMemoryContextProvider against either backend without changing a line of the memory logic itself.
A deterministic smoke test
Every memory system needs one test that answers a single question without ambiguity: does it actually recall the specific fact it was told. I originally wrote this test by asking the agent a question in a second session and asserting on the model’s final text output, and it was flaky, not because the memory system was broken, but because the model would sometimes phrase the answer in a way my string match didn’t anticipate. The fix was to stop testing the model’s prose and start testing the retrieval layer directly, which is the part that’s actually deterministic.
using Xunit;
public class MemoryRecallSmokeTests
{
[Fact]
public async Task Recalls_CloudProvider_Fact_Across_Sessions()
{
var store = new SqliteMemoryStore(dbPath: ":memory:can-share-connection-in-real-setup");
var userId = $"smoke-test-{Guid.NewGuid()}";
// Session 1: seed the fact the way StoreAIContextAsync would have,
// via a canned extraction result rather than a live model call, so
// this test has no dependency on sampling variance.
var extracted = new OrgMemoryState { CloudProvider = "Azure", Region = "westeurope" };
var existing = await store.LoadStateAsync(userId, default) ?? new OrgMemoryState();
await store.SaveStateAsync(userId, new OrgMemoryState
{
CloudProvider = extracted.CloudProvider ?? existing.CloudProvider,
Region = extracted.Region ?? existing.Region,
DeploymentStrategy = existing.DeploymentStrategy,
AdditionalFacts = existing.AdditionalFacts,
}, default);
// Session 2: a brand new provider instance, same user ID, nothing
// shared except what's in the store. This is the actual assertion:
// does the provider surface the fact on a cold session. Call the
// real public entry point, InvokingAsync, not the protected
// override, so this test exercises exactly what the agent runtime
// would call. FakeAgent/FakeSession are thin test doubles with no
// behavior beyond satisfying the constructor, omitted here for space.
var provider = new OrgMemoryContextProvider(store, extractionClient: null!, userId);
var invokingContext = new InvokingContext(new FakeAgent(), new FakeSession(), new AIContext());
var aiContext = await provider.InvokingAsync(invokingContext);
var summary = aiContext.Messages!.Single().Text;
Assert.Contains("Azure", summary);
Assert.Contains("westeurope", summary);
Assert.DoesNotContain("AWS", summary);
}
}
The part worth calling out: this test never calls a real language model. It seeds the store the way a successful extraction would have, then verifies that a cold ProvideAIContextAsync call, from a completely new provider instance, surfaces exactly the fact that was written. That's the honest boundary of what "deterministic" can mean for a system with a model somewhere in its pipeline: you make the retrieval and merge logic deterministic and test that rigorously, and you test extraction quality separately, with looser, statistical evaluation, because extraction is the one place an LLM's judgment is actually load-bearing.
Production considerations, the honest section
Auth mapping. Never let a client hand you the userId it wants memory keyed on. Map the session's authenticated principal, whatever your auth middleware already validated, to a stable internal user ID server-side, and treat any user-supplied identity claim in a request body as untrusted input. The failure mode if you get this wrong isn't subtle: one user's session ends up reading another user's memory, because your context provider trusted a field it should have derived instead.
Retention and consent. Extracted facts are still personal or organizational data, and “we cached it in a memory store” doesn’t exempt it from whatever data handling rules already apply to the rest of your system. Concretely: put a TTL on raw session turns (they’re the least necessary thing to keep and the most sensitive, since they’re closer to verbatim conversation), keep extracted state around longer but make it inspectable, and build a real forget-me path, a method that deletes a user’s row from userMemory entirely rather than soft-marking it, if your retention policy calls for hard deletion. If you're storing anything a person might reasonably not want remembered indefinitely, ask once, up front, rather than assuming silence is consent.
Catching a wrong memory. This is the one I underestimated. An extraction step can misread a hedge as a commitment, “we might move to multi-cloud eventually” is not the same claim as “we are moving to multi-cloud,” and once that’s written into userMemory it will confidently color every future turn until someone notices. Three things helped: logging the source turn alongside every extracted fact, not just the fact itself, so a human reviewing a wrong memory can see exactly what triggered it; building a boring internal page that lists a user's current memory state in plain text, because "can I just look at what it thinks it knows" turned out to be the fastest way anyone on the team actually caught a bad extraction; and treating any correction the same way you'd treat a genuine update, overwrite the field and keep a record that it was corrected, rather than pretending the wrong version was never there.
So where do I actually draw the line
Here’s my answer, and it’s a size threshold, not a philosophy.
If you’re on a single repository small enough that one or two engineers hold the real conventions in their heads, and the code you’d grep is genuinely current, I’m with Theo. Skip the memory layer. Every fact you cache outside the code is a fact that can drift from the code, and on a small enough team someone will notice the drift before it matters. Build the context provider above if you want, but point it at nothing more than session-scoped scratch state, and don’t pretend you need userMemory at all yet.
The line I’d personally use is somewhere around the point where a codebase has more than one area nobody on the current team fully owns, the kind of area where the last person who understood why a decision was made has moved teams or left, or where “the standard” is something new hires get told in a hallway conversation rather than something the repository states anywhere a grep would find it. That’s usually somewhere past ten engineers and past a single repository, which lines up with Ahuja’s “first ten people” instinct more than I expected it to when I started this. Below that line, curated memory is solving a problem you don’t have yet. Above it, the code stops being a reliable narrator of its own history, and the one sentence of memory that tells an agent “we are an Azure shop, this is a standard, not a preference” is the difference between a correct first attempt and a plausible-looking wrong one that costs someone an afternoon to catch.
Neither camp is wrong. They’re describing different codebases.
Tags: dotnet, ai-agents, agentic-memory, azure-cosmos-db, csharp, llm-engineering, software-architecture
Top comments (0)