We were pairing on a Thursday. Nothing dramatic. A junior on my team, maybe eight months in, sharp, the kind who actually reads the code instead of pasting it. We'd asked the AI for a chunk of logic, it came back in a few seconds, clean and typed and already passing the tests he'd written.
He went to accept it. I said, "no, don't ship that."
He stopped. Looked at it again. Looked at me. And he asked the most reasonable question in the world.
"Wait — how did you know that's wrong?"
And I opened my mouth to explain, and nothing came out.
I knew. I just couldn't say how.
Here's the thing that rattled me, because it wasn't a dramatic bug and this isn't a dramatic story. The diff was wrong. I was right that it was wrong, and ten minutes later we proved it. That part was fine.
What wasn't fine was that I had no idea how to tell him what I'd done. I hadn't run a checklist. I hadn't spotted a specific line and matched it to a rule. Something about the shape of it had made the back of my neck go cold, and I'd learned, over a lot of years, to trust that exact feeling. But "trust the cold feeling on the back of your neck" is not a thing you can hand to a human being who hasn't grown the neck yet.
I tried anyway. I said something about how it felt off, how the error path looked too convenient, how I'd want to know what happens if this gets called twice. All true. All useless to him, because every one of those was a conclusion, and he was asking for the method, and I didn't have a method. I had a scar that fired before I'd finished reading.
That's the part I've been chewing on since. Not that I couldn't teach him. That the single most valuable thing I do all day is the one thing I have no idea how to transmit.
Two things get bundled into the word "senior"
For years I thought being senior was a pile of knowledge. And some of it is, and that part I can teach. "Validate at the boundary." "Don't trust the client's timestamp." "Wrap the external call." Rules. I can write them on a whiteboard and he'll have them by Friday, and honestly, the AI already has all of them and applies them more consistently than I do.
But the thing I did on Thursday wasn't a rule. It was the opposite of a rule. It was knowing, against a clean diff that satisfied every rule either of us could name, that something was still wrong. Rules tell you what to check. The other thing tells you to keep looking after every check has passed. One is knowledge. The other is judgment, and I have never once been able to put it into words that survive contact with someone who hasn't earned it.
You don't learn judgment. You survive into it — and that's exactly why I couldn't hand it to him.
Nobody taught me the thing I did on Thursday. There was no course, no cert, no senior who sat me down. I got it the only way I think anyone gets it: I was confidently wrong about something that mattered, in front of someone who paid for it, and the lesson got welded on.
Where the cold feeling actually came from
Let me be specific, because I can trace that exact flinch to its source.
Years ago I shipped a write path that told the client "got it" before it had actually saved the row. The code was clean. Genuinely clean — idiomatic, typed, tests around it, error handling buttoned up. It read like someone careful wrote it, because I was careful. By every rule I knew at the time, it was correct.
Then one ordinary day a retry landed at exactly the wrong moment. The "got it" went out, the save never happened, and a paying customer got locked out of their own account with nothing in the logs to say they'd ever been there. Ack before persist. I can still feel the phone call.
That's the source. When I looked at Thursday's diff and my neck went cold, I wasn't running an algorithm. I was pattern-matching against that customer. The AI's clean code looked exactly as correct as my clean code had looked, right up until it cost someone their account. The flinch is that customer, compressed into half a second and welded onto my nervous system by the worst night of that year.
So when the junior asked "how did you know," the honest, full answer was: a person I locked out at 2am four years ago taught me, and I have no idea how to give you that without the person. You can't lecture someone into the flinch. You can only get unlucky enough times that it grows on its own.
And here's the part that actually scares me
For twenty years, the flinch had a reliable supply chain. You got it by doing the grind — writing the production code yourself, shipping it, and being there when it broke. The small failures stacked up into judgment whether you wanted them to or not. You couldn't skip the grind, so you couldn't skip the scars.
Watch what we just did to that supply chain.
The junior doesn't write the production code anymore. The AI does, and it does it well, and I'm glad — the typing was never the hard part. But the typing, the shipping, the breaking, the 2am call: that was the entire mechanism that used to turn a junior into me. We didn't just automate the boring part. We automated the forge. I got my judgment because I had to do, by hand, the exact work the AI now does for him before he ever feels it go wrong. He gets the clean output on day one and never has to earn the neck.
So the one skill that survived AI — the only one, the "no" — is also the one skill the AI quietly stopped manufacturing in the next generation. The machine can hand him the code. It cannot hand him the scar, and the scar was the teacher.
To be clear, this is not "make juniors suffer"
I'm not romanticising pain, and I'm not saying pull the AI and make him hand-write CRUD until he bleeds for it. That's cargo-cult mentorship — the suffering was never the point, the consequence was. And I use AI all day; I'd never go back. The problem isn't that he has it easy. The problem is narrower and weirder than that: we removed the thing that used to produce judgment, and we haven't replaced it with anything.
So that's the actual job now, and it's a design problem, not a vibes problem. I can't give him the flinch. But I can give him the one thing that grew the flinch in me — being wrong somewhere it's cheap to be wrong — on purpose, faster, instead of waiting for a real customer to do it the expensive way.
A few things I actually do with him now:
I make him the skeptic, not the author. The AI writes the diff. His job isn't to write a better one — it's to break the one we got. Before we accept anything, he has to finish the sentence "this loses money when ___." He's wrong most of the time. Doesn't matter. The rep is the distrust, not the catch. That's the muscle, and it only grows when it's his job to doubt.
I manufacture the consequence small. Ship it behind a flag, to one internal user, to a canary. Let it break where breaking is cheap, so the lesson arrives before the customer does. A flinch you got from a staging incident is the same flinch — it just didn't cost anybody their account.
I stop answering "is it right" and make him answer "how is it wrong." The question trains the instinct. "Looks good" trains nothing. If I just tell him what I saw, he learns my conclusion. If I make him hunt for how it breaks, he starts growing his own neck.
Why this is the exact reason I build the way I do
One level up, same problem.
I work on an agent platform, and the whole industry right now wants to cheer for the thing that produces. Look how much it ships, look how clean. But a thing that produces is just doing the skill that stopped being scarce — and worse, it's doing it in a way that signs off on its own work, the same confident "looks correct" over a masterpiece and a disaster alike. It's the AI version of a junior with no flinch: fast, fluent, and completely unable to distrust itself.
So I never let the thing that writes the code be the thing that blesses it. There's an author that produces the diff, cheap and endless. There's a separate skeptic whose entire job is to distrust that diff and try to break it — the flinch given its own seat, institutionalised, so it doesn't depend on anyone having gotten unlucky enough to grow one. And there's a human on the merge button, because somebody still has to own the call, and that human is the one still accumulating the real scars. Author, skeptic, human. That's the whole shape of xenition, and it's my actual answer to the junior's question: you can't teach the flinch, so you build the seat that does its job and you put the person in it until they grow their own.
I still can't tell him how I knew. What I can do is put him where he'll find out the way I did — wrong, early, and somewhere cheap enough that the lesson costs a canary instead of a customer. The knowing was never teachable. The getting-wrong is. That's the only part I can actually hand him.
Top comments (7)
you can teach someone the rules, but a lot of judgment comes from knowing what actually broke before.
that’s one reason i like having production history in the review loop. PRI (Production Reliability Index) gives the PR another signal based on things like fragility, churn and past incidents when that data is connected.
yeah, that's a sharp connection — PRI is basically trying to do with data what I spent four years growing in my nervous system the hard way.
the thing that gives me pause, in the best way: fragility and incident history is the closest thing we have to a transmissible scar. I couldn't hand the junior the 2am phone call, but "this file has broken three times in prod and churns every sprint" — that I can put on the screen. it won't give him the flinch, but it points his attention at the exact places mine fires, before he's earned it himself.
where I'd be careful is the same place I'm careful with the AI: a signal that's good enough can quietly replace the judgment instead of training it. if the PR comes back green on PRI and he reads that as "safe," we've just built a more sophisticated version of "tests pass, ship it." the whole point of the flinch is that it keeps looking after every signal clears.
so I'd want PRI as a reason to look harder at a diff, never as a reason to look less. frame it as "history says this neighborhood is dangerous — now you go find how this change breaks," and it's doing exactly what I'm trying to do with him by hand. that's the useful version.
would genuinely like to hear how you're surfacing it in the loop — is it a score on the PR, or does it actually redirect the reviewer's attention to specific lines?
nothing replaces what four years of 2am calls taught your nervous system, and i wouldn't claim a score does. what it can do is hand a small, rough copy of that history to someone who hasn't lived it yet.
to your question at PRI it's a score first, and it shows up in three places.
in the editor, the free VS Code / Cursor plugin shows PRI while you're writing. it's read only, never modifies files, and only needs a sign in, no API key. so the junior sees the score before a PR exists, which is probably the most useful moment for training.
on the PR, the GitHub App reads the changed files and posts a comment with the overall PRI (0–100), plus fragility, governance compliance and code volatility, with the reasoning and evidence behind the verdict. 70 is the ready line. permissions are contents read, PR read/write for the comment, and email read. it never commits or branches, and needs no config file or CI change.
in the agent: there's an MCP server for Claude Code, Claude Desktop, Cursor or any streamable HTTP client, so the agent writing the code can pull the same score. full scans there can take up to 30 minutes. it also runs alongside CodeRabbit over MCP if you already use that for line review.
underneath, PRI rolls up seven sub indices: fragility, drift, governance compliance, runtime signals, code volatility, deployment velocity and escalation. runtime signals and escalation need observability and ticketing data connected, and until then they show "No data" instead of guessing. scores also calibrate to each org over time.
on line level: as far as i know it doesn't mark specific lines yet, i'll check and come back.
your green means safe worry is the right one. the gate is advisory by default, teams decide which policies become blocking, a human can always override, and confidence is capped at 85 so it never claims certainty. a green score without history behind it is a much weaker green.
so i'd frame it the way you did. a red score tells you where to start looking. a green one only means nothing in the available signals flagged it.
The ten minutes after "don't ship that" are the part I'd hand the junior, even if the flinch itself can't be handed over. The flinch tells you where to look; proving it is a method, and a method is something he can watch and copy. Your write-path story also already turned into one of the hints you gave him: "what happens if this gets called twice" is a scar packed into a question, and the question works for someone who doesn't have the scar yet. Did the test you wrote to prove it end up in the suite?
You caught the thing I underweighted, and you're right. I wrote the whole post around the part that can't be handed over and skipped past the part that can. The flinch tells you where to look — that's the unteachable half. But the ten minutes after, the actually-proving-it, is a method, and he can watch a method and copy it. More than that: once I've said "don't ship that," the suspicion is already free. He doesn't need to have grown the neck to run the proof — he just needs someone to point. So the division of labor is cleaner than I made it sound: I supply the flinch, he practices the method, and the method is his to keep.
And "what happens if it's called twice" — yeah, that's the move I didn't notice I was making. That's a scar compressed into a question, and the question survives the handoff even though the scar doesn't. A library of those is the closest thing to a transferable flinch that exists. The one caution I'd put on it, because I've watched it happen: a stack of scar-questions calcifies into a checklist, and a checklist does the exact thing the flinch is supposed to defend against — it makes him feel safe the moment the list passes. The questions are training wheels for the neck, not a substitute. The day the bug is shaped like nothing I packed into a question, he's back to needing the thing I can't give him. Which is fine — that's just the next scar, and now he has the method to prove it when it comes.
On your actual question: yes, it went in the suite, and that turned out to be the quiet lesson. A scar written down as a regression test is the flinch given a permanent seat — it fires forever without anyone having to remember the 2am call. That's the same "institutionalize the skeptic" move as everything else I do, just at the smallest scale. But here's the limit that keeps it honest: the test protects us against that bug for good, and does exactly nothing for the one neither of us has met yet. The suite is a graveyard of past flinches. The neck is for the graves we haven't dug. So — the test is in there, and the junior's the one who wrote it, which was the point.
Which production responsibility would you leave with the junior so a bad decision still has a visible owner?
good question, and it's the one that actually keeps me honest — because it's easy to say "human on the merge button" and quietly mean "senior on the merge button." if the junior never owns anything that can hurt, he never grows the neck. he just watches me have one.
so the thing I leave with him is the merge itself, on his own surface area, with his name on it. not "assist the senior" — own the call. he's the one who clicks it, he's the one in the channel when the canary goes sideways, he's the one who writes the "here's what I missed" after. the author (AI) wrote the code, the skeptic (also him, mostly) tried to break it, but the decision to ship is his and it's attributable to him.
what makes that survivable instead of reckless is the blast radius, not the supervision. I don't shrink his ownership, I shrink the crater. flag it, canary it, one internal cohort — so when he's wrong, and he will be, the owner is clearly him and the cost is clearly a rollback instead of a customer. that's the whole trick: real ownership, fake stakes, for as long as it takes the stakes to stop being fake on their own.
the moment I can't afford to let him own a call is exactly the moment he most needs to have owned twenty smaller ones already. if the first thing he owns is the scary thing, I waited too long.