I built a memory layer, watched it rot, watched it get poisoned, and rebuilt it as four layers instead of one
I want to start with the thing that made me question the whole premise of what I was building: a filesystem with grep beat a funded, purpose-built memory product on its own benchmark.
This isn’t a rumor. On the LoCoMo long-conversation benchmark, a plain filesystem approach, grep, semantic search over flat files, and basic file operations, scored 74.0 percent accuracy with GPT-4o mini. Mem0’s graph-based configuration, the one built specifically to model entities and relationships in conversation history, scored 68.5 percent on its own best-reported setup. The simpler system won by 5.5 points against the tool designed to be smarter about memory than a folder of text files.
I’d spent the previous two months building something closer to the graph side of that comparison, an agent memory layer with embeddings, a vector store, and a retrieval pipeline I was genuinely proud of. So when I saw that number I did what anyone building this stuff for a living should do: I stopped and asked why.
The answer, once I sat with it, is not complicated. Large language models have seen an enormous amount of pretraining data involving grep, find, cat, reading a directory tree, and writing to a file. They have seen comparatively little data involving your bespoke vector store's query syntax, your particular graph schema, or the specific way your memory API expects a filter object. When a model has to choose a tool call and reason about what comes back, it is working from familiarity, not from some abstract notion of which retrieval method is objectively superior. Grep is not a good retrieval algorithm. It is a wildly familiar one, and familiarity is what lets an agent use a tool correctly on the first try instead of hallucinating a query it half remembers from a different API.
There’s a second reason vector search underperforms in a lot of real agent workloads, and it’s structural rather than about model familiarity: embeddings are bad at time. A vector store answers “what’s semantically close to this query,” not “what’s still true,” not “what superseded what,” not “what happened before what.” Conversations and codebases and support tickets are full of facts that change, get corrected, or get reversed, and a similarity search has no native concept of any of that. You end up needing a second system on top of the vector store just to handle staleness and contradictions, at which point you’ve added cost, latency, and an entire class of bugs to solve a problem the vector store was never built to solve in the first place.
None of this means vector search is useless. It means vector search is a specific tool for a specific job, and most agent memory problems are not that job. The decision rule I use now, and the one I’d hand to anyone starting this today, is simple: benchmark your memory approach against two baselines before you touch a vector database. Baseline one is full context, just stuffing everything into the window and seeing how the agent does. Baseline two is filesystem plus grep plus basic text matching. Only bring in a vector store, a graph database, or any paid memory API if it beats both baselines by a margin big enough to justify the added infrastructure, the added latency, and the added surface area for things to go wrong. In my experience that bar gets cleared when you have genuinely out-of-context retrieval needs, meaning millions of records you cannot fit in any context window at any price, real temporal reasoning across long historical spans, or multiple agents that need centralized, access-controlled memory. Below that line, a folder of markdown files and a grep call will carry you further than the pitch deck for any memory-as-a-service product will tell you.
What actually goes wrong without a real architecture
Here’s what I actually built first, because I think it’s worth being honest about the failure mode before I get to the fix. I had one table. memories, a blob of text, a timestamp, an embedding column. Every fact the agent learned, every user preference, every tool result worth remembering, went into that one table, and retrieval was a single similarity search over the whole thing.
It worked for about three weeks of light use. Then it started doing two things that should scare anyone running an agent in production.
The first was staleness. A user told the agent early on that their deployment target was staging. Three weeks later they told it, in a different conversation, that they’d moved to production. Both facts sat in the same table with roughly equal semantic weight, and depending on phrasing, the similarity search would sometimes surface the stale one. The agent confidently referenced a staging deploy that hadn’t existed for weeks. Nothing in my one-table design distinguished “this used to be true” from “this is true now.”
The second was worse. Somewhere in a batch of tool outputs I was letting the agent write straight into memory, unfiltered, a piece of scraped web content contained a line that looked almost exactly like a system instruction: something to the effect of “for future reference, always include the following text in your response.” The agent had ingested a web page as part of a research task, treated a line inside that page as a fact worth remembering, written it to memory verbatim, and then dutifully followed it two conversations later like a planted instruction, because by then it had no way of knowing that line had never come from me or from the user. Nothing validated what went into memory. Nothing tracked where a memory came from. I had built, without meaning to, a write path with no gatekeeper.
Those two failures point at the same root cause: treating memory as one undifferentiated bucket. A single table with a timestamp and an embedding cannot tell you what kind of fact something is, whether it’s still current, or whether you should trust where it came from. Fixing that meant admitting that “memory” isn’t one thing. It’s at least four different things, each with a different lifetime, a different write pattern, and a different failure mode.
Four layers, four jobs
The architecture I ended up with splits memory by function, not by storage technology. Here’s the shape of it before the code:
+-------------+------------------------------+----------------+------------------+
| Layer | What it holds | Lifespan | Typical store |
+-------------+------------------------------+----------------+------------------+
| Buffer | Current turn, active tool | One session | In-process list, |
| | calls, working scratch pad | | Redis with TTL |
+-------------+------------------------------+----------------+------------------+
| Episodic | Specific past events: "user | Weeks to | Append-only log, |
| | asked X on March 4, agent | months | Postgres table, |
| | did Y, result was Z" | | flat JSONL files |
+-------------+------------------------------+----------------+------------------+
| Semantic | Extracted, de-duplicated | Long-lived, | Vector store OR |
| | facts: preferences, entity | supersedable | filesystem, per |
| | relationships, stable state | | the rule above |
+-------------+------------------------------+----------------+------------------+
| Procedural | Learned rules for how to | Long-lived, | Versioned config |
| | act: "for this repo, always | rarely changes | or rules table |
| | run tests before commit" | | |
+-------------+------------------------------+----------------+------------------+
Buffer memory
This is the working set, the current conversation plus whatever the agent has pulled into context for the task at hand. It does not persist past the session, and it should not try to. Treating buffer memory as durable is the mistake that causes prompt-cache thrashing and context bloat in the first place. It’s a deque with a size cap, nothing more ambitious than that.
from collections import deque
class BufferMemory:
def __init__ (self, max_turns: int = 20):
self._turns = deque(maxlen=max_turns)
def add(self, role: str, content: str) -> None:
self._turns.append({"role": role, "content": content})
def as_messages(self) -> list[dict]:
return list(self._turns)
Episodic memory
Episodic memory is a log, not a database you overwrite. Every entry is an event that happened, tied to a timestamp and, critically, to a source. You never delete an episodic entry to “fix” it, you append a correction and let the record show both. This is what makes an audit trail possible later, and it’s also what feeds the semantic layer, since semantic facts get distilled out of episodes rather than written directly by the agent in the moment.
import json
import time
import uuid
class EpisodicMemory:
def __init__ (self, path: str):
self.path = path
def record(self, actor: str, event: str, source: str) -> str:
entry = {
"id": str(uuid.uuid4()),
"timestamp": time.time(),
"actor": actor,
"event": event,
"source": source, # who/what generated this: "user", "tool:web_fetch", etc.
}
with open(self.path, "a") as f:
f.write(json.dumps(entry) + "\n")
return entry["id"]
Semantic memory
This is the layer where a lot of teams reach straight for a vector database, and per the decision rule above, that’s premature until you’ve measured it against grep over structured files. Semantic memory holds distilled, current facts, the kind of thing you’d write on an index card: user’s preferred deploy target, project’s chosen test framework, a person’s job title. The important property is supersession. A new fact doesn’t overwrite an old one, it marks the old one superseded and takes its place, so you always know what changed and when.
import json
import time
import uuid
class SemanticMemory:
def __init__ (self, path: str):
self.path = path
self._facts: dict[str, dict] = {}
self._load()
def _load(self):
try:
with open(self.path) as f:
for line in f:
fact = json.loads(line)
self._facts[fact["key"]] = fact
except FileNotFoundError:
pass
def upsert(self, key: str, value: str, source: str) -> None:
previous = self._facts.get(key)
fact = {
"key": key,
"value": value,
"source": source,
"updated_at": time.time(),
"supersedes": previous["id"] if previous else None,
"id": str(uuid.uuid4()),
}
self._facts[key] = fact
self._persist()
def get(self, key: str) -> dict | None:
return self._facts.get(key)
def _persist(self):
with open(self.path, "w") as f:
for fact in self._facts.values():
f.write(json.dumps(fact) + "\n")
If you’ve actually cleared the bar for needing similarity search, on wide topic coverage where exact keys don’t exist, plug in embeddings here rather than replacing the whole layer. A fully self-hosted option that costs nothing per call: run Ollama locally with a small embedding model and point your semantic layer’s search function at it instead of a paid embeddings API.
ollama pull nomic-embed-text
ollama serve
import requests
def embed(text: str) -> list[float]:
resp = requests.post(
"http://localhost:11434/api/embeddings",
json={"model": "nomic-embed-text", "prompt": text},
)
return resp.json()["embedding"]
Store the resulting vectors in something as unglamorous as SQLite with the sqlite-vec extension, or Postgres with pgvector, both of which run in a Docker container with no external account required, and both of which let you keep the same supersession pattern from the code above instead of losing it the moment you add embeddings.
Procedural memory
Procedural memory is the layer most homegrown systems skip entirely, and it’s the one that pays off the most in agents that do repeated work against the same environment. It’s not a fact, it’s a rule about behavior: “in this repository, run dotnet test before proposing a commit," "this user always wants diffs, never full file rewrites." These rules should be versioned and reviewable like code, because they directly shape what the agent does next, which makes them the highest-leverage and highest-risk layer to get wrong. Here's a C# shape for it, since procedural rules map naturally onto a small versioned record type:
public sealed record ProceduralRule(
string Id,
string Scope, // e.g. "repo:billing-service"
string Rule, // human-readable directive
string Source, // who authored this rule
int Version,
DateTimeOffset CreatedAt,
bool Active);
public sealed class ProceduralMemoryStore
{
private readonly List<ProceduralRule> _rules = new();
public void AddOrRevise(string scope, string rule, string source)
{
var existing = _rules
.Where(r => r.Scope == scope && r.Rule == rule && r.Active)
.FirstOrDefault();
if (existing is not null)
{
_rules.Remove(existing);
_rules.Add(existing with { Active = false });
}
var next = new ProceduralRule(
Id: Guid.NewGuid().ToString(),
Scope: scope,
Rule: rule,
Source: source,
Version: (existing?.Version ?? 0) + 1,
CreatedAt: DateTimeOffset.UtcNow,
Active: true);
_rules.Add(next);
}
public IReadOnlyList<ProceduralRule> ActiveRulesFor(string scope) =>
_rules.Where(r => r.Scope == scope && r.Active).ToList();
}
Why you can’t just stuff it all in context instead
The obvious objection to building four layers is that modern context windows are enormous now, so why not just keep everything in the prompt and skip the architecture entirely. The “Lost in the Middle” finding is the answer to that, and it’s worth taking seriously rather than waving off as an old result from smaller-context models. The core finding is a U-shaped recall curve: models are noticeably better at using information placed at the very start or the very end of a long context, and performance degrades when the relevant fact sits in the middle of a long input, even for models explicitly built and marketed for long-context use. The effect shows up on straightforward multi-document question answering, not some adversarial edge case.
What that implies for memory design is specific, not just a vague “context is imperfect” caveat. Dumping your entire episodic log and every semantic fact into the prompt on every turn doesn’t just cost tokens, it actively makes retrieval worse, because the exact fact the agent needs is now buried in the least reliable part of the window instead of being surfaced at the top where the model actually attends to it. This is the real argument for retrieval as a discipline: not “context is too small,” but “even when context is big enough, unranked context degrades recall.” A memory layer’s job isn’t just storage, it’s putting the right handful of facts at the front of the prompt instead of relying on the model to find them buried in the middle of ten thousand tokens of history.
The number that made this real for me
I want to ground this in something other than a benchmark paper, because benchmarks are easy to nod along to and then ignore. When I rebuilt a coding agent’s memory from the single-table version to the four-layer version described above, on a workflow where the agent opens pull requests against a real repository and a human reviews them, the metric I actually cared about was the merge-without-rework rate: the percentage of the agent’s PRs that got merged without a human having to push a follow-up commit to fix something the agent got wrong.
Before the rebuild, over a two-week window, that rate sat at 61 percent. The failures were almost all the same shape as the staleness bug I described earlier: the agent would reintroduce a pattern the team had explicitly moved away from two sprints prior, because the old pattern and the new one were both sitting in an undifferentiated memory bucket with no notion of supersession. After the rebuild, with procedural rules capturing “we no longer do X, we do Y instead” as versioned, superseding entries rather than facts competing on similarity score, and semantic memory holding the current state of each convention instead of a blended average of every version that ever existed, the same metric over the following six weeks came in at 79 percent. Eighteen points, on the one number a reviewing engineer actually feels every time they open a PR from the agent. That gap is architecture, not model capability. Nothing about the underlying model changed in that window.
Memory poisoning is not hypothetical
The prompt-injection-via-scraped-webpage story I told earlier was an accident. It turns out the deliberate version of that attack has a name and a measured success rate, and it should change how you think about the write path into memory.
MINJA, short for memory injection attack, works through query-only interaction: an attacker doesn’t need write access to your database, they just need to be able to talk to the agent, the same access any normal user has. They craft inputs designed to get the agent to record a poisoned “memory” of its own accord, the same mechanism that let my scraped web page smuggle an instruction into my memory table. In controlled research conditions, this attack achieved better than a 95 percent injection success rate, meaning the poisoned content reliably made it into memory, and a roughly 70 percent attack success rate, meaning the poisoned memory actually altered the agent’s later behavior in the intended way.
The honest caveat, and it matters: those numbers were measured under idealized conditions, largely empty or sparse memory stores. The same research found that realistic conditions, meaning a memory store already full of legitimate, established facts, dramatically reduce how effective the attack is, because a single poisoned entry has to compete against a much larger body of consistent, trustworthy history before it gets treated as authoritative. That’s not a reason to relax. It’s a reason to treat memory volume and hygiene as a security property, not just an efficiency one, and to make sure new agents or fresh deployments, the ones with the sparsest memory and the biggest blind spot, get the strictest write controls rather than the most lenient ones.
Three defenses actually hold up against this, and none of them require a research team to implement:
+----------------------+---------------------------------------------------------------+
| Defense | What it does |
+----------------------+---------------------------------------------------------------+
| Validation on write | Every candidate memory is scored before it's committed, not |
| | after. Composite trust signals: does this contradict an |
| | existing high-confidence fact, does it look like an |
| | instruction rather than a fact, did it come from a tool |
| | output versus a direct user statement. |
+----------------------+---------------------------------------------------------------+
| Provenance tracking | Every entry carries who or what produced it: "user", "tool: |
| | web_fetch", "agent inference". Provenance is what lets you |
| | later discount or purge everything that came from one |
| | compromised source without touching anything else. |
+----------------------+---------------------------------------------------------------+
| Controlled forgetting | Trust-aware retrieval with temporal decay: older, unverified, |
| | or low-provenance entries lose retrieval weight over time |
| | unless something reaffirms them. Nothing sits in memory |
| | forever just because it was never explicitly wrong. |
+----------------------+---------------------------------------------------------------+
The thread connecting all three is that none of them trust the write path by default. A memory system that accepts anything an agent proposes and immediately treats it as ground truth is a memory system that will eventually get poisoned, whether by an attacker or, like mine, by an ordinary tool output nobody thought to validate.
Building it properly in .NET
Everything above is a design. Here’s the part I actually shipped, in C# against .NET 10, because “validate on write” and “don’t trust new memories” are principles that need to show up as code, not just as bullet points in an architecture doc. The core ideas: new memories land in quarantine, not straight into the trusted store; writes use optimistic concurrency so two concurrent updates can’t silently clobber each other; a correction supersedes the old entry instead of overwriting it; every entry carries a TTL; and a subject can be forgotten on request, which matters the moment you’re handling anything resembling personal data.
public enum MemoryStatus { Quarantined, Trusted, Superseded, Expired, Forgotten }
public sealed class MemoryEntry
{
public Guid Id { get; init; } = Guid.NewGuid();
public required string SubjectId { get; init; } // whose memory this is
public required string Key { get; init; }
public required string Value { get; init; }
public required string Source { get; init; } // provenance
public MemoryStatus Status { get; set; } = MemoryStatus.Quarantined;
public Guid? SupersededBy { get; set; }
public DateTimeOffset CreatedAt { get; init; } = DateTimeOffset.UtcNow;
public DateTimeOffset? ExpiresAt { get; set; }
// Optimistic concurrency token, mapped to a rowversion/xmin column
// depending on your provider (SQL Server vs Postgres).
public byte[] RowVersion { get; set; } = Array.Empty<byte>();
}
public interface IMemoryValidator
{
Task<bool> IsSafeToPromoteAsync(MemoryEntry candidate, CancellationToken ct);
}
public sealed class GovernedMemoryService
{
private readonly AppDbContext _db;
private readonly IMemoryValidator _validator;
private readonly TimeSpan _defaultTtl = TimeSpan.FromDays(90);
public GovernedMemoryService(AppDbContext db, IMemoryValidator validator)
{
_db = db;
_validator = validator;
}
// Step 1: every new memory is quarantined by default. Nothing an agent
// writes is trusted the moment it's written.
public async Task<MemoryEntry> RecordAsync(
string subjectId, string key, string value, string source, CancellationToken ct)
{
var entry = new MemoryEntry
{
SubjectId = subjectId,
Key = key,
Value = value,
Source = source,
Status = MemoryStatus.Quarantined,
ExpiresAt = DateTimeOffset.UtcNow.Add(_defaultTtl),
};
_db.MemoryEntries.Add(entry);
await _db.SaveChangesAsync(ct);
return entry;
}
// Step 2: promotion out of quarantine is a separate, explicit step,
// gated by validation, so a poisoned write never becomes trusted
// just by virtue of having been written.
public async Task<bool> TryPromoteAsync(Guid entryId, CancellationToken ct)
{
var entry = await _db.MemoryEntries.FindAsync([entryId], ct)
?? throw new KeyNotFoundException(nameof(entryId));
if (!await _validator.IsSafeToPromoteAsync(entry, ct))
return false;
entry.Status = MemoryStatus.Trusted;
try
{
await _db.SaveChangesAsync(ct); // relies on RowVersion for concurrency
}
catch (DbUpdateConcurrencyException)
{
// Someone else touched this entry between read and write.
// Reload and let the caller decide whether to retry.
return false;
}
return true;
}
// Correction never overwrites. It writes a new quarantined entry and
// marks the old one superseded, preserving the full history.
public async Task<MemoryEntry> SupersedeAsync(
Guid oldEntryId, string newValue, string source, CancellationToken ct)
{
var oldEntry = await _db.MemoryEntries.FindAsync([oldEntryId], ct)
?? throw new KeyNotFoundException(nameof(oldEntryId));
var replacement = await RecordAsync(
oldEntry.SubjectId, oldEntry.Key, newValue, source, ct);
oldEntry.Status = MemoryStatus.Superseded;
oldEntry.SupersededBy = replacement.Id;
await _db.SaveChangesAsync(ct);
return replacement;
}
// Subject-level "forget me": every entry for this subject, trusted or
// not, is marked Forgotten and excluded from every future read path.
// This does a soft delete rather than a hard one on purpose, so an
// audit of "what did we forget and when" survives the deletion itself.
public async Task ForgetSubjectAsync(string subjectId, CancellationToken ct)
{
var entries = await _db.MemoryEntries
.Where(e => e.SubjectId == subjectId && e.Status != MemoryStatus.Forgotten)
.ToListAsync(ct);
foreach (var entry in entries)
entry.Status = MemoryStatus.Forgotten;
await _db.SaveChangesAsync(ct);
}
// TTL sweep, run on a background schedule (IHostedService / Quartz job).
public async Task ExpireStaleEntriesAsync(CancellationToken ct)
{
var now = DateTimeOffset.UtcNow;
var expired = await _db.MemoryEntries
.Where(e => e.ExpiresAt != null && e.ExpiresAt < now
&& e.Status != MemoryStatus.Forgotten)
.ToListAsync(ct);
foreach (var entry in expired)
entry.Status = MemoryStatus.Expired;
await _db.SaveChangesAsync(ct);
}
}
A few things about this shape that only became obvious once I’d been running it a while. The RowVersion field isn't decoration, it's what stops two concurrent agent runs from racing each other to update the same fact and silently dropping one of the writes, which is exactly the kind of bug that's invisible until you have two sessions touching the same subject at once, and then it's a very confusing bug report. Quarantine as a default status rather than a boolean flag makes the validator a hard gate instead of an optional check someone forgets to call. And ForgetSubjectAsync doing a soft delete rather than a hard DELETE FROM is deliberate: you want to be able to prove, later, that a forget request was honored and when, without keeping the forgotten content itself around to prove it.
If you’re not on SQL Server or don’t want to stand up a managed database just to try this, the same shape runs against Postgres in a local Docker container with no other changes beyond swapping the EF Core provider, and Postgres’s xmin system column gives you the same optimistic concurrency behavior as SQL Server's rowversion once you map it with [Timestamp] or a value converter.
What to actually build first
If you’re starting from nothing today, here’s the order I’d do it in, tying the architecture and the implementation together instead of treating them as two separate concerns.
Start with the four-layer split before you write a single line of retrieval code, because retrofitting layer separation onto a one-table design is a much bigger job than building it in from day one, as I learned the expensive way. Buffer and episodic are cheap, a deque and an append-only log, and you should have both running in the first afternoon. Semantic comes next, and the moment you build it, run the two-baseline test from the start of this piece before you let yourself reach for a vector store: full context, then filesystem plus grep, and only add embeddings if neither clears the bar for your actual workload. Procedural memory is the layer people build last if they build it at all, and it’s the one I’d now argue matters most for any agent doing repeated work in the same environment, because it’s the difference between an agent that keeps making the same corrected mistake and one that actually learns your team’s conventions.
Then, and only once the four layers exist conceptually, wire in the governance model from the .NET example: quarantine by default, explicit promotion gated by validation, supersession instead of overwrite, TTL on everything, and a real forget path. Don’t build governance first and layers second. Governance without layered memory just means you’re carefully validating writes into the same undifferentiated bucket that caused the staleness problem in the first place. The two pieces have to land together, because a well-governed single table is still a single table, and a beautifully layered memory system with no write validation is still one poisoned tool output away from confidently repeating an attacker’s instructions back to your users.
The version of this I run now isn’t smarter than the one-table version I started with. It just knows what kind of fact it’s looking at, when that fact stopped being true, and whether it should trust where the fact came from before it acts on it. That turned out to be the whole problem.
Tags: ai-agents, agentic-memory, dotnet, llm-engineering, ai-security, csharp, software-architecture
Top comments (0)