DEV Community

Cover image for My support bot stopped forgetting everything when I changed AI memory write triggers
Lars Winstand
Lars Winstand

Posted on Originally published at standardcompute.com

My support bot stopped forgetting everything when I changed AI memory write triggers

I had a support bot that looked great in demos and acted like it had amnesia in production.

The failure mode was embarrassingly simple:

  • customer already rebooted the router three times
  • customer already confirmed the firmware version
  • customer already asked for email follow-up instead of SMS
  • bot asked them to do all of that again

The stack was not exotic:

  • n8n for orchestration
  • pgvector for long-term memory
  • an OpenAI-compatible API for model calls
  • GPT-5.4 for most runs
  • Claude Opus 4.6 for harder replies

The model had the transcript. The model had a big context window. The workflow had retrieval.

And it still behaved like it had never seen the case before.

The fix was not a better model.

The fix was changing when memory gets written.

The rule that fixed my support agent: write memory only on 5 event types — confirmed facts, preference changes, state-changing troubleshooting steps, decisions, and outcomes.

Once I did that, retrieval got smaller, cleaner, and actually useful.

The real bug: I stored chat history instead of case history

My first version wrote almost every turn into pgvector.

That included junk like:

  • "yes"
  • "still broken"
  • "I already tried that"
  • "ok"
  • "can you explain that again"

At first this feels safe. More data should mean better recall, right?

I think that is exactly backwards for support workflows.

For support bots, full-transcript memory is usually worse than event memory because it:

  • pollutes retrieval
  • increases repeated mistakes
  • bloats prompts
  • burns tokens reloading low-value context

If your workflow runs continuously in n8n, Make, Zapier, OpenClaw, or a custom agent loop, this gets worse over time. Every ticket adds more clutter. Every retrieval step gets noisier. Every prompt carries more baggage.

The bad fix I tried first

Naturally, I tried the wrong fix before the right one.

I changed the memory writer to summarize every conversation chunk before storing it.

That sounded smarter. It was not.

Now I had compressed noise instead of raw noise.

GPT-5.4 and Grok 4.20 both did the same thing overloaded memory systems usually do: they latched onto plausible but low-value details, then repeated steps that were already resolved.

This is the part a lot of people miss when building support automations:

Bigger memory is not better memory.

If your memory layer is sloppy, the model does not become more informed. It becomes more distractible.

The 5 memory write triggers that actually worked

I stopped writing every turn and only wrote memory on 5 event types.

1. Confirmed facts

Store facts only when they are explicitly confirmed.

Good examples:

  • Customer is on Netgear Nighthawk R7000
  • Account is on the Pro plan
  • ISP is Comcast
  • Device is running firmware 1.0.11

Bad examples:

  • Customer probably has an old router
  • This sounds like a DNS issue
  • User may be on the free tier

If the model inferred it, it does not belong in long-term memory.

2. Preference changes

These are easy to ignore and they matter constantly.

Examples:

  • Customer prefers email over phone
  • Do not restart devices during business hours
  • Contact only after 5 PM local time
  • Wants concise replies, not step-by-step explanations

A support bot can be technically correct and still feel broken if it forgets preferences.

3. State-changing troubleshooting steps

This was the biggest one.

Store the step only if it changed system state.

Examples:

  • Factory reset completed
  • DNS changed from ISP default to Cloudflare
  • Firmware rolled back from 1.0.11 to 1.0.9
  • Router moved from bridge mode to router mode

Do not store suggestions.

Bad memory:

  • Suggested factory reset
  • Recommended checking DNS
  • Asked customer to reboot

That is not state. That is conversation.

4. Decisions

Examples:

  • Escalate to Tier 2
  • Pause automation and wait for screenshot
  • Schedule on-site technician visit
  • Close loop pending customer response

Decisions shape what the next run should do. They belong in memory.

5. Outcomes

Examples:

  • Issue resolved after firmware rollback
  • Ticket closed unresolved due to inactivity
  • Packet loss persisted after modem replacement
  • Escalation accepted by network team

Outcomes stop the agent from reopening dead paths.

The rule in table form

Event Type Store It?
Confirmed hardware model Yes
Customer says "ok" No
DNS actually changed Yes
Bot suggested DNS change No
Customer changed contact preference Yes
Conversation summary of small talk No
Escalate to Tier 2 Yes
Issue resolved after rollback Yes

What the memory payload started to look like

Instead of stuffing transcript chunks into pgvector, I started storing structured events.

Something like this:

{
  "ticket_id": "T-18422",
  "event_type": "state_change",
  "timestamp": "2026-02-14T16:22:11Z",
  "summary": "Factory reset completed on Netgear Nighthawk R7000",
  "metadata": {
    "device": "Netgear Nighthawk R7000",
    "action": "factory_reset",
    "status": "completed"
  }
}
Enter fullscreen mode Exit fullscreen mode

Or this:

{
  "ticket_id": "T-18422",
  "event_type": "preference_change",
  "timestamp": "2026-02-14T16:31:02Z",
  "summary": "Customer prefers email follow-up instead of SMS",
  "metadata": {
    "channel": "email"
  }
}
Enter fullscreen mode Exit fullscreen mode

That gave retrieval something useful to rank.

A practical write filter

This is the kind of logic I wish I had started with.

export function shouldWriteMemory(event: {
  type: string;
  confirmed?: boolean;
  changedState?: boolean;
}): boolean {
  switch (event.type) {
    case "confirmed_fact":
      return !!event.confirmed;
    case "preference_change":
      return true;
    case "state_change":
      return !!event.changedState;
    case "decision":
      return true;
    case "outcome":
      return true;
    default:
      return false;
  }
}
Enter fullscreen mode Exit fullscreen mode

And the retrieval side stayed aggressive about keeping context small.

export function buildMemoryContext(memories: MemoryRecord[]): string {
  return memories
    .sort((a, b) => b.score - a.score)
    .slice(0, 5)
    .map((m) => `- [${m.event_type}] ${m.summary}`)
    .join("\n");
}
Enter fullscreen mode Exit fullscreen mode

That slice(0, 5) mattered more than I expected.

Not 50 records. Not the whole ticket history. Just the highest-signal events.

How this changes the prompt

Before:

Here is the previous conversation history:
[huge blob of transcript chunks and summaries]
Enter fullscreen mode Exit fullscreen mode

After:

Relevant case memory:
- [confirmed_fact] Customer is on Netgear Nighthawk R7000
- [state_change] Factory reset completed
- [state_change] DNS changed from ISP default to Cloudflare
- [decision] Escalation to Tier 2 pending screenshot
- [preference_change] Customer prefers email follow-up

Do not repeat completed troubleshooting steps.
Continue from the current case state.
Enter fullscreen mode Exit fullscreen mode

That second prompt is shorter and much harder for the model to misunderstand.

What changed in production

The behavior shift was immediate.

The bot stopped:

  • asking customers to repeat resolved steps
  • re-suggesting fixes that already failed
  • losing handoff context between runs
  • wasting tokens on giant memory payloads

It got better at:

  • continuing open tickets correctly
  • handling asynchronous follow-ups
  • escalating with cleaner summaries
  • respecting customer preferences

This also improved reliability in long-running workflows. If your agent is running all day in n8n or Make, bad memory rules compound. Good memory rules compound too.

Why this matters even more if you pay per token

This is not just a quality problem. It is also a cost problem.

Bad retrieval design quietly becomes an expensive habit:

  • larger prompts
  • more retrieval overhead
  • more retries after wrong answers
  • more model calls because the workflow keeps looping on already-resolved issues

If you are paying per token, noisy memory punishes you twice: once in quality and again on the bill.

This is exactly why I like using an OpenAI-compatible layer that does not make me obsess over every token while I tune workflows.

With Standard Compute, I can keep the same OpenAI SDK patterns, route across models like GPT-5.4, Claude Opus 4.6, and Grok 4.20, and run agent-heavy automations without watching a per-token meter all day.

That does not remove the need for good memory design. Nothing removes that. But it does make iteration a lot less annoying when you are testing retrieval, prompts, and multi-step automations in n8n or custom agent loops.

If you are building a support bot, my advice is simple

Treat memory like a timeline of state changes, not a backup of the conversation.

Support is not a creative writing task. It is closer to a state machine with human language wrapped around it.

Store:

  • what is true
  • what changed
  • what was decided
  • what happened next

Do not store every sentence just because you can.

That one change fixed more of my support bot's "memory" than switching models ever did.

If your agent keeps forgetting obvious things, I would look at memory write triggers before I blame the model.

Top comments (0)