It's been about two months since I started actually running this, not just building it, and I keep coming back to the same question instead of moving past it.
What I'm building
I work at Wontopos, a memory API in the same space as Mem0 and Zep. Store what a user says once, retrieve only what's relevant on each query, hand that to an LLM prompt. The long term goal, and I'll admit this out loud, is to become something close to standard infrastructure for AI memory at the B2B level. Not a feature bolted onto someone else's product, but the layer everyone building on AI just reaches for.
That's a long way off, and I'm aware of how that sounds coming from two months in. Right now what actually matters is much smaller. Get real developers to use it, and find out where we're wrong before we've committed to being wrong at scale.
Where the confidence went
This is the part I didn't expect. Watching actual usage come in has made me less confident, not more. I thought two months of real traffic would sharpen my sense of what people want from something like this. Instead it mostly sharpened my sense of what I don't know. People reach for memory for reasons I didn't anticipate, and skip it for reasons I also didn't anticipate. Neither of those is something I can fix by thinking about it harder alone.
What I'm actually asking
So here's the honest question, and I mean it plainly rather than as a lead-in to a pitch.
Where would you actually reach for a memory API. And what would you need it to do that you're not getting from whatever you use now, whether that's another memory product, rolling your own with a vector store, or just stuffing more into the context window and hoping.
Even a short answer helps. I'd rather hear "I wouldn't use this and here's why" than nothing at all, because silence doesn't tell me whether we're solving a real problem or a problem we invented.
Top comments (6)
A real answer from someone who rolled their own: I'm building Mythex (an AI app builder), and we built memory in-house rather than using a memory API. Storage and retrieval weren't the slow part. These were:
So what would make me reach for an API: that write policy and provenance built in, a list I can render for users with edit and delete, and per-user delete and export in one call. At our size those mattered more than retrieval quality.
That's basically a checklist for what a memory layer has to get right beyond retrieval. Thanks for writing it all out, I'm saving this.
woochan, from the side of building LLM integrations for clients, here's where I'd actually reach for a memory API and what I'd need.
The clearest case is a support or sales assistant that talks to the same customer over weeks. Right now we either re-read the whole conversation history into the context window (expensive, and it degrades past a few turns) or we hand-roll a vector store and a retrieval prompt that we keep tuning. A memory API that stores "what the user said once" and hands back only the relevant slice would replace that hand-rolled layer, and the thing I'd need it to do that I'm not getting from a raw vector store is forgetting and updating: when a user changes their mind, the old fact should stop being retrieved, not just sit there with a lower score.
The second case is a personal assistant or a long-running agent where the context window is the bottleneck. Here the value is that retrieval is scoped to the right entity (this user, this project) automatically, rather than me writing that scoping logic every time.
Two honest caveats, since you asked for them: for short-lived, single-session apps a memory API feels like overkill, and I'd just stuff the context. And the thing that would make me trust it over rolling my own is a clear, inspectable "why was this retrieved" — otherwise I can't debug a wrong answer.
Alex, small web & automation studio
Appreciate you spelling out the use cases and being honest about where a memory API is overkill. That's more useful to me than a polite yes. Thanks for taking the time.
Glad it was useful! Saw in your bio that you're building an open benchmark for agent memory — that's exactly what I'd want before picking any memory layer. When it's public, I'd be happy to run it against our own setup (BM25 + embeddings over a task knowledge base) and share the numbers back.
Alex
Wrote a full rundown of what this actually does a couple weeks ago, linking it here instead of the post so this doesn't turn into a pitch: An Introduction to Wontopos.Knowing what I'm talking about will probably get you to a more useful answer than guessing at it.