Nobody touched the code. Nobody changed the model. I just tidied my mailbox. And my classifier lost an entire class, without a single warning.
TL;DR: my email triage bot classifies with a TF-IDF model trained on my own folders. A manual re-sort drained the Personnel class from 196 examples to 9. The next retraining pushed it under the minimum-class threshold and deployed a 6-class model instead of 7. Silently. Here is the incident, and the three guardrails I built so it can never happen again.
This article is for everyone who can retrain a model with one command, and has no alarm around it.
The setup
My email triage bot has been running for months. I told its story in Put the LLM Last: rules first, a classifier next, the LLM as a last resort.
The classifier is TF-IDF with logistic regression. Training happens in Python, inference in pure Go. Parity between both was verified: 5,895 emails out of 5,895, same verdict.
Labels come from real usage: the folder each email ends up in. 5,895 labeled emails, 7 classes kept for training.
And retraining takes one command: make retrain. That convenience is what bit me.
The incident: nobody touched the code
One weekend, I cleaned up my mailbox. The emails in the Personnel folder moved to other, better-organized folders.
On the data side: the Personnel class went from 196 examples to 9.
The training pipeline has a healthy threshold: a class with too few examples is excluded from the model. Healthy, except the exclusion was silent.
So the next make retrain trained a 6-class model instead of 7. It got deployed. No message, no refusal, no visible difference.
The symptom came later: personal emails filed anywhere. The pipeline did exactly what it was told. That is precisely the problem.
Why it is structural: your labels are alive
This bug is not a local blunder. It is structural, and it is waiting for you too.
My labels are generated by a living system: my mailbox and my sorting habits. One re-sort, one archive, one new habit, and the class distribution shifts.
Your dataset is not a file. It is a snapshot of something that moves.
A retraining pipeline only sees today's snapshot. Without a memory of the previous model, it cannot know a class disappeared. You have to give it that memory.
Guardrail 1: expected classes
First guardrail, the simplest. The retraining script first reads the deployed model's classes. Then it requires finding them in the data.
# retrain.sh reads the deployed model's classes,
# then enforces them on the new training run
train.py --expect-labels Cabinet,Maison,Newsletter,...
# an expected class under the minimum threshold:
# ABORT. Nothing written, nothing deployed.
A class can still disappear. But it is now an explicit human decision, not a side effect of tidying up.
Guardrail 2: a quality gate against the deployed model
The first guardrail catches a vanishing class. It misses a class that survives but collapses.
So every model now embeds its metrics in its own file: accuracy, macro-F1, per-class F1. Macro-F1 averages each class's F1: small classes weigh as much as big ones.
On the next retraining, the script reads the deployed model's macro-F1 and takes it as the floor. The candidate regresses beyond the tolerated margin? Abort, before anything is written.
The gate is tested like code: a simulated collapse must trigger the abort. An untested guardrail is a decoration.
Guardrail 3: the escalation monitor
The first two guardrails protect training. The third one watches production.
When the classifier is unsure, it escalates: below a confidence threshold, the decision goes to the fallback LLM. That escalation rate is the model's heartbeat.
Every decision writes one journal line: timestamp, confidence, escalated or not, label. Never the subject or the sender, the journal is persistent.
A small tool aggregates by week and alerts when the recent rate exceeds the baseline by more than 10 points. When the world changes, the model doubts more. Doubt is measurable.
What I did not do
I did not rebuild the Personnel class. The data was scattered and polluted, and rebuilding would have cost more than the class was worth.
Personnel went back to deterministic rules, with the LLM as a net. The 6-class model became the documented, accepted reference.
Not every class deserves ML. A class poor in examples lives better as rules.
The checklist before your next retrain
Three guardrails, one evening of work. The incident costs weeks of trust.
- [ ] Make the script list the deployed model's classes before training
- [ ] Refuse the retrain if an expected class falls under the threshold
- [ ] Embed the metrics inside the deployed model's file
- [ ] Compare candidate to deployed: regression beyond the margin, abort
- [ ] Test your guardrails: a simulated collapse must trigger the abort
- [ ] Journal every production decision: confidence, escalation, no content
- [ ] Alert on the escalation rate, not only on errors
- [ ] Accept moving a class out of ML when rules do better
What to remember
A retraining pipeline without memory destroys your model silently, because data moves without notice.
The defense has three parts: expected classes, a quality floor read from the deployed model, and a watched escalation rate.
None of this requires an MLOps platform. An honest script is enough.
An ML pipeline to make trustworthy, even a small one? Let's talk.
Sources: Google, Rules of Machine Learning (monitoring and degradation) · scikit-learn, f1_score and macro averaging · Put the LLM Last, the bot's architecture
Top comments (0)