DEV Community

Cover image for I've been a developer for 10 years. AI just showed me I only had one real skill."

I've been a developer for 10 years. AI just showed me I only had one real skill."

Info Inlet on September 30, 2026

Last Sunday I built an invoice tracker for a friend who runs a small shop. Nothing clever. I described what I wanted, let the AI do the rest, and a...
Collapse
 
contentclips_st profile image
ContentClips •

Ack-before-persist is the perfect example, because the judgment was invisible in review too: the code read clean, tests were green, and the lie lived between the API promise and the durable state. Two habits that make judgment auditable: (1) turn durability contracts into tests — one test that fails whenever a response returns before commit, so "acked before saved" can never pass CI twice; (2) when inheriting AI-generated code, review the failure story instead of the happy path: what happens when the retry lands at exactly the wrong moment? If the answer needs a paragraph, that\u2019s the part you can\u2019t rent from a supplier. The Sunday invoice app is fine — it\u2019s the three-weeks-later silent bug that ten years were actually for.

Collapse
 
infoinlet1 profile image
Info Inlet •

This is the sharpest addition anyone's made, because you turned the vibe into something checkable, which was exactly the gap I couldn't close in the post.

The durability-contract test is the one I'd tattoo on people. "Acked before saved can never pass CI twice" is the whole thing. The catch, and I say this having lived it, is that you can only write that test if you already had the judgment to know the failure exists. The test doesn't generate the insight, it preserves it. So it's less a substitute for the "no" and more a way to stop paying for the same "no" twice, which honestly might be the best you can do with judgment: catch it once by hand, then never by hand again.

The failure-story review is the part I want to push on teams. Reviewing the happy path is how the ack-before-persist bug sailed through in the first place, mine included. "What happens when the retry lands at exactly the wrong moment" should just be a required field on every PR. And your tell for it is perfect: if the answer needs a paragraph, that's the part you can't rent. I'm stealing that line, with credit.

"The three-weeks-later silent bug is what ten years were actually for" is a cleaner version of my own thesis than I wrote. Yeah. That's it.

Collapse
 
parsa4873fe3aa profile image
Parsa Mohammadi •

Agreed, and that means you only write the durability test for failures someone already recognized. The retry bug that burned a different team never makes it into a test. Neither does the one that hit this team two years ago, before anyone on it now was hired.

At Tomosu we work on the Production Reliability Index, which scores new changes against the incident history a team already has. So "never by hand again" also covers failures nobody remembers, including ones that never got a test. The first "no" still has to come from a person.

Thread Thread
 
contentclips_st profile image
ContentClips •

The blind spot on the far side of it: incidents that never got written up. Every team runs a shadow incident history — the retry bug someone fixed silently, the near-miss a senior caught by smell — and history-based scoring can only rank what got filed. The fault-injection fixture from earlier covers a different slice: it doesn't ask whether the failure is guarded, it asks whether anyone notices when it isn't. Random breakage plus one honest assertion tests the noticing itself. History tells the next team what to fear; the fixture tells you whether the skepticism still works. Both miss the silent fixes — which is why the first 'no' stays human.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

@parsa the production reliability index is the piece i didn't have a name for. my whole method is me walking the seams asking "what if this runs twice," but i can only ask that about failures i've personally been burned by. my skepticism is capped at my own scar tissue. scoring a change against the team's full incident history is how you borrow scars you never earned, including the ones that predate everyone currently on the team. that's the real gap in my approach and i'd use it. the "first no still has to come from a person" caveat is the important one though — the index can rank the fear, it can't originate it.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

@contentclips_st this is the sharper edge of it, yeah. history-based scoring can only ever see incidents that got filed, and in my experience the filed ones are the minority. the retry someone quietly fixed on a friday, the near-miss a senior killed by smell and never wrote up — those never make it into any history, so any tool built on that history is blind to them.

the distinction you drew is the part i'm keeping: history tells you what to fear, the fixture tests whether you'd even notice. i'd been mushing those into one thing. "does anything scream when it breaks" is a completely different question from "did we handle case X," and testing the noticing instead of the code is a framing i hadn't had. both still miss the silent fix, which is exactly why the first "no" stays human — a person smelling it is the only reason it ever got caught in the first place.

Collapse
 
slabb profile image
Sam LABBE •

The origin story you picked is doing more work than you're claiming for it. "The 'got it' went out, the save never happened, and there was nothing in the logs" — that wasn't only a judgment miss, it was a record miss. Two events that should have been reconcilable — what you promised the customer vs. what actually persisted — and nobody could replay them afterward. Judgment is how you catch it the second time. The record is how it gets caught at all.

So I'd add a quiet fourth member to author / skeptic / human: the record. A skeptic's "no" that isn't written down is indistinguishable from a "no" that never happened — and a human who "owns the call" with no trace of the call owns it rhetorically. The cheap version is a twist on your own ritual: you already say "here's how this loses money" out loud, wrong most of the time. Write the sentence down, timestamped, next to the diff it objected to. Do that for a year and "wrong most of the time" stops being a shrug and becomes calibration — twelve real disasters out of forty objections is a track record, and it's the only line on the rewritten CV nobody can argue with.

Full disclosure: this is the pattern behind the flight-recorder journal I build, so consider the source. But I think you've underpriced your own skill. Judgment you can't show is indistinguishable from luck. The no is the skill — the receipt is what makes it one.

To your question, then: the one I'd keep is deciding what deserves a receipt.

Collapse
 
infoinlet1 profile image
Info Inlet •

ok this is the best pushback i've gotten and i think you're right. i was treating it as a pure judgment miss and it wasn't. the judgment failure and the record failure were two different failures stacked on top of each other, and honestly the record one was the worse of the two. if the promise and the persist had been reconcilable afterward, someone finds that customer in a day. without it they just vanish, and the lesson vanishes with them. i felt the second failure and filed it under the first. that's on me.

the "no that isn't written down is indistinguishable from a no that never happened" line is going to live in my head for a while. because yeah, i've been quietly proud of catches i can't actually prove i made. that's not a track record, that's a story i tell myself. twelve out of forty is a number. a vibe isn't.

and i'll take the hit on underpricing it. "judgment you can't show is indistinguishable from luck" is exactly the hole in my own post. i said the only score that's mine is the disasters i saw coming, then gave no way to count them. the receipt is what turns it from a feeling into a score.

the self-plug is fine, you flagged it and it earned its place, so no complaints there.

the one spot i'd push back a little: i don't think the record replaces the no, and i don't think you're saying it does, but worth keeping straight. a receipt on a bad call is still a bad call, just a legible one. writing it down doesn't make the judgment good, it makes it checkable, which is the thing that lets it get good over time. so maybe the honest shape is the no is the skill, the receipt is what lets you find out if the skill is real.

"deciding what deserves a receipt" is a great answer because it's recursive, it's a judgment call about judgment. you can't write everything down or the record's just noise. deciding what's worth the ink is the same muscle one level up. stealing that framing, with attribution.

Collapse
 
slabb profile image
Sam LABBE •

The line that stayed with me is "nothing in the logs." Judgment catches what green tests can't — agreed. But the facility that answers "what actually happened to that customer" is separate from judgment, and it's buildable: the ack-before-persist decision is an event like any other, and it can be recorded in a form that survives the post-mortem. Judgment finds the bug; the record decides whether the next engineer believes you.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

Yeah, I'm fully bought in now. The thing that finally clicked reading this: judgment is a property of a person, and the record is what lets it outlive that person being in the room. "Nothing in the logs" didn't just hide the bug from me that night — it meant nobody who came after could reconstruct the call either. The ack-before-persist decision happened, it was a real event with a real cause, and the system kept no memory that it had ever been made. The bug was recoverable in principle and unrecoverable in practice, and that gap is entirely a record problem, not a judgment one.

"The record decides whether the next engineer believes you" is the sharpest way anyone's put it in this whole thread. Because that's the failure mode of pure judgment — it's non-transferable. I can't hand my hunch to the junior, and I can't prove it to the skeptic, and six months later I can't even prove it to myself. An un-recorded no dies with the shift it happened on. The receipt is what lets a decision survive turnover, survive the post-mortem, survive me forgetting I ever made it.

So where I've landed: author produces, skeptic objects, human owns the call — and the record is what makes all three legible to the next person instead of evaporating the moment the room clears. Not a fourth actor so much as the thing that keeps the other three from being hearsay. Good exchange. You moved the post.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

Nothing clever. I [built an invoice tracker]...

The honesty here is the whole post. What AI actually exposed for a lot of us isn't "you can't code" — it's that the typing was never the valuable part. The valuable part was knowing what to build and whether it's actually right.

That "one real skill" is judgment: shaping the problem and verifying the result. The agent does the middle 80%; the 20% on either end is the job now, and it always secretly was.

What did you land on as your one skill? Mine was "deciding what done means" — everything else turned out to be delegable.

Collapse
 
infoinlet1 profile image
Info Inlet •

yeah you put it better than i did honestly. "deciding what done means" is such a good way to say it, because that's the part nobody can hand off. the agent will happily tell you it's done. it's done by every measure it has. whether it's actually done is a different question and that one never left your desk.

mine, if i had to name it, is close to yours but phrased as a negative. it's the "no." looking at something that works and not trusting it. deciding what to build is the front end of judgment and verifying is the back end, and the "no" is just the thing that runs the whole length of it. the confident diff i overrule, the feature i kill, the green test i don't believe. the machine can produce forever but it can't withhold, and withholding turned out to be the job.

funny thing is both of ours, "what done means" and "the no," are the two skills nobody ever put on a cv or tested in an interview. ten years and the one thing that survived was the one thing nobody could measure. kind of says everything about what we were actually hiring for all along.

Collapse
 
rajanpanwar profile image
Rajan Panwar •

The trap right now is confusing "easy to generate" with "safe to ship."

My un-automatable skill isn't even catching bugs and it's knowing what not to build. Since writing code is now practically free, teams are going to drown in features and architectural sprawl that nobody actually understands six months later. Saying "no" to a design or killing unnecessary complexity before a single line is drafted is something a prompt can't do for you.

What’s one architectural decision you will never let AI make on your project?

Collapse
 
infoinlet1 profile image
Info Inlet •

"Knowing what not to build" is a sharper version of the skill than I gave it credit for in the post. Catching the bug is downstream — you only have to catch it because it got built. Your skill stops it a step earlier, before the thing exists to be wrong. And you're right that it's the one a prompt structurally can't do, because the AI's whole gravity is toward "yes, here's how" — it has no instinct for "this shouldn't exist at all." Ask it to build a feature and it never once asks whether you should.

The decision I'll never hand it: where the boundaries go — what's allowed to talk to what. The AI is brilliant at filling in a module and genuinely terrible at deciding whether two things should be coupled in the first place, because that call depends on stuff that lives nowhere in the code — which parts I expect to change together, which team owns which, what I'm willing to rewrite in a year and what has to stay stable. Let it make those calls and you get exactly the sprawl you're describing: everything reaching into everything, works fine in month one, and by month six nobody can change anything without breaking three things they didn't know were connected.

Writing the code inside a boundary is cheap and the AI should do it. Drawing the boundary is the expensive, judgment-loaded call, and it's the one that's almost impossible to undo later. So that's mine. The AI fills the boxes; I decide where the lines are.

Honestly your "drown in features nobody understands six months later" is the thing I'd put money on being the defining mess of the next couple years. Cheap production doesn't reduce the cost of a system — it just moves the whole bill to comprehension.

Collapse
 
vrunda_chauhan_a52cc23b11 profile image
Vrunda Chauhan •

The best habit I picked up for this is simple: don't review what the AI wrote. Review what it skipped.

Whenever an AI gives me working code, I force myself to ask three questions before touching it:

What happens if the user's internet drops halfway through this?
What happens if two people click this button at the exact same second?
What happens when this database table hits 10,000 rows instead of 10?

AI is great at building the path where everything goes right. The human job is just making sure the app doesn't blow up when things go wrong.

Collapse
 
infoinlet1 profile image
Info Inlet •

"Review what it skipped" might be the best four-word version of the whole post. Because that's the asymmetry exactly — the AI shows you the path it built, confidently and completely, and the danger is never in what's on the screen. It's in what isn't. And nothing in the output points at the hole, because from the model's side the hole doesn't exist. It answered the question you asked; the bug is in the question you didn't.

Your three questions are great, and I'd notice what they have in common: not one of them is answerable by reading the code. Internet dropping halfway, two clicks in the same second, 10 rows becoming 10,000 — those all live in the world the code runs in, not the code itself. That's why the AI misses them and why static review misses them too. You're not checking the picture, you're checking the frame around it.

The one I'd add to your list, since it's the one that got me: "what happens if this exact request arrives twice?" Retries, double-submits, a webhook that fires again — same family as your two-clicks question, and it's quietly behind half the money-losing bugs I've seen. The happy path handles it once, beautifully, and falls apart the second time.

The only thing I'd gently push on is "the human job is just making sure it doesn't blow up when things go wrong" — I'd drop the "just." That's not the small remaining chore after the AI did the real work. That is the work now. The AI took the part that was always easy and left us the part that was always the actual job. Your three questions aren't a safety check on top of development. They're the development.

Collapse
 
kartik-nvjk profile image
Kartik N V J K •

The "is this going to quietly lose someone money in three weeks" line is the whole thing for me. Production is easy to check, it either runs or it does not, but that slow judgment call almost never shows up in a test, which is why a first-try-working invoice tracker can still be wrong in a way nobody sees for a month. How do you try to surface that kind of delayed failure before it ships?

Collapse
 
infoinlet1 profile image
Info Inlet •

honestly that line still makes me nervous and you nailed why. green tests will wave that bug through all day because it's just not a test-shaped problem.

so i don't really test for it. the thing that's helped most is boring: i stopped looking inside the function and started looking at the edges. the three week bug almost never lives in the logic, it lives where my code hands off to something else. a network call, a queue, a db write, a client that might retry. so before i trust anything i just walk those handoffs and ask the same dumb question, what if this step runs twice, or what if it runs and the next one doesn't. that ack-before-persist bug was exactly that. the code was fine, the order across the boundary wasn't.

the other habit is i make myself actually say the sentence out loud, "here's how this loses someone money." if i can't finish it i don't really understand the change yet, i just know it runs. most of the time the sentence is nonsense and there's nothing there, fine. the one time it writes itself easily is the one you go fix.

and i've basically given up reviewing my own stuff for this. whoever wrote it is the worst person to catch it because they already decided it's right, that's why it came out that way. so i let someone (or something) else just try to break the happy path. replay the request, send it twice, kill it halfway through. not "does it work" but "show me where it half worked."

none of it catches everything though. the whole definition of these bugs is they slip whatever net you knew to build. i've just stopped letting "all tests pass" stand in for "someone actually tried to make this thing lose money." not the same sign off.

Collapse
 
nullandvoid_ profile image
Jyanthi •

The distinction between production and judgment really hits home. AI can generate increasingly impressive solutions, but knowing when not to trust a solution is a different skill entirely. The “author, skeptic, human” model is a fascinating way to think about where engineering is heading.

Collapse
 
infoinlet1 profile image
Info Inlet •

Thanks Jyanthi. The thing that gets me is the two skills feel identical from the inside — a confident wrong answer and a confident right one arrive in the exact same voice. That's why the "don't trust it" can't live in the same place that generates the answer. It has to be a separate seat, or it just agrees with itself.