DEV Community

Cover image for I've been a developer for 10 years. AI just showed me I only had one real skill."
Info Inlet
Info Inlet

Posted on

I've been a developer for 10 years. AI just showed me I only had one real skill."

Splitting coding into production versus judgment

Last Sunday I built an invoice tracker for a friend who runs a small shop. Nothing clever. I described what I wanted, let the AI do the rest, and about forty minutes later it was running. Routes, the database schema, login, even the fiddly error handling I used to be a bit smug about. Worked on the first try.

I should have just felt good about it, and mostly I did. But there was another feeling underneath that I couldn't shake for the rest of the evening. That thing had done, in forty minutes, the stuff I spent most of my twenties learning how to do.

Not "the robots are taking our jobs." Something more specific, and more personal than that. I looked at ten years of work I was proud of and realised I couldn't tell anymore how much of it had ever really been mine.

Quick disclaimer before anyone heads for the comments: I use AI to write code every day now and I'm not giving that up. This isn't that kind of post. But that Sunday made me ask something I'd genuinely never had a reason to ask, because until recently you couldn't take the job apart to look at it. Of everything I call "my skills," which one was actually mine?

Turns out "being a developer" was two jobs, not one

For about ten years I never had to separate them, because you couldn't get good at the second one without grinding through the first.

Call the first one production. Making the thing exist. The language, whatever framework everyone used that year, the APIs you eventually know by heart, wiring one piece to another until an idea turns into something that runs. This is the part bootcamps teach and interviews test. It's what my CV was, basically top to bottom. And it's the part AI is now flat-out better at than me. Which, honestly, fine. It was never the hard part. It just took years to get quick at.

The second one I don't really have a clean word for. Judgment, maybe. It's looking at something that works and knowing, somehow, that it's wrong. Not "does it run" but "is this going to quietly lose someone money in three weeks." Knowing which confident, tidy, all-tests-passing answer you shouldn't trust. Nobody ever put this on my CV. There was no course for it, no cert, no exam. It felt less like a skill and more like a hunch you can't explain to a junior no matter how you try.

AI took the first job completely. And because the two were always tangled together, it feels like it took "being a developer." It didn't. It took production, which was most of what I trained, and it hasn't laid a finger on the judgment, because the judgment was never in the code. It was in me.

So a decade gave me one skill, not ten

This is the part that actually kept me up.

Every skill I'd have bragged about, anything I'd have listed if you asked why I was worth hiring, was production. Recall. The framework, the right flag, the pattern, the clean CRUD. All of it gone in an afternoon, and done well. Ten years of stacking it up, handed back to everyone for free in a single product update.

Ten years didn't build me ten skills. It built one. The other nineteen I was renting from a supplier who just took them back.

The one thing that didn't get taken was the judgment. And here's what gets me: I'd never once counted it as a skill. It felt like taste. It felt like something I was, not something I'd learned. Which is exactly why no course sells it and no interview tests it. You can't really teach the "no."

Everything I was proud of, AI does in an afternoon now. The one thing it can't do was the only thing that was ever mine.

The bug that made it obvious

Let me get specific, because I paid for this one.

Years back I shipped a write path that told the client "got it" before it had actually saved the row. The code was clean. Genuinely clean. It read like someone careful wrote it, because I did. Everything from those ten years was in there: idiomatic, properly typed, tests around it, error handling buttoned up. By every measure I'd have used at the time, some of the best code I'd written.

Then, one ordinary day, a retry landed at exactly the wrong moment. The "got it" went out, the save never happened, and a paying customer got locked out of their own account with nothing in the logs to say they'd ever been there. Ack before persist.

Here's what stuck with me. None of my ten years caught that. The syntax was perfect. The framework was used correctly. The tests were green. Every one of those rented skills did its job. The only thing that would have saved that customer was the skill I'd never even named: the reflex to look at working code and not trust it, to stop and ask "okay, what happens if this gets retried," to say no to my own clean diff.

The code scored an A on everything AI now beats me at. It failed on the one thing that was mine.

I'm not mourning the other nineteen, to be clear

I'm not about to start hand-writing CRUD to prove a decade meant something. For production, AI is a better developer than I am at 5pm on a Friday, and leaning on it is just correct. Take the scaffold. Take the fix. Ship faster. That part isn't the loss.

The trap is quieter than that. When a tool does all nineteen of your rented skills, the easy conclusion is that you've been made pointless, when the truth is closer to the opposite. It cleared away everything that was never really the point and left you standing on the one thing that was. And if you can't tell which skill was the real one, you'll spend the next ten years defending the nineteen it already took, instead of getting better at the one it can't.

So, a few things I actually do now.

I rewrite the CV around the "no." Not "expert in [framework]" β€” the machine is an expert in that too. The line worth writing is the judgment call: the bug I didn't trust when it looked fine, the design I killed, the confident diff I overruled and turned out to be right about.

I keep the muscle warm on purpose. Before I accept anything the AI hands me, I make myself say one sentence out loud: "here's how this loses money." I'm wrong most of the time. Doesn't matter. The rep is the distrust, not the catch.

And I stopped measuring myself in output. Lines, tickets, features shipped. Those are the machine's numbers now, and racing it on them is just volunteering to be a slower model. The only score that's still mine is the disasters I saw coming that it didn't.

And this is more or less why I build the way I do

One level up, same logic.

I work on an agent platform, and the whole industry right now wants to cheer for the thing that produces. Look how much it ships. But a thing that produces is just doing the skill that stopped being scarce. Endless production is endless output that nobody has actually vouched for. The scarce thing was never the code. It was someone looking at the code and saying no, which is the one skill a decade left me with.

So I don't let the thing that writes the code be the thing that signs off on it. There's an author that produces, all nineteen rented skills, cheap and endless. There's a separate skeptic whose entire job is to distrust the output and try to break it, the "no" given its own seat. And there's a human who owns the call, because that was always the only part that was ours. Author, skeptic, human. That's the whole shape of xenition, and it's the same thing the locked-out customer taught me. Producing the thing was always cheap. Knowing it was wrong was the job.

Ten years didn't make me a great producer. AI just proved I never had to be. What it made me, slowly and expensively and one disaster at a time, is someone who can look at something that works and tell it's wrong. That's the one real skill, and it's the only one AI handed back.

Honest question for the comments: of everything you know, which single skill would still be yours if AI could do all the rest tomorrow? The one that isn't syntax, isn't a framework, isn't just recall. πŸ‘‡

Top comments (24)

Collapse
 
contentclips_st profile image
ContentClips •

Ack-before-persist is the perfect example, because the judgment was invisible in review too: the code read clean, tests were green, and the lie lived between the API promise and the durable state. Two habits that make judgment auditable: (1) turn durability contracts into tests β€” one test that fails whenever a response returns before commit, so "acked before saved" can never pass CI twice; (2) when inheriting AI-generated code, review the failure story instead of the happy path: what happens when the retry lands at exactly the wrong moment? If the answer needs a paragraph, that\u2019s the part you can\u2019t rent from a supplier. The Sunday invoice app is fine β€” it\u2019s the three-weeks-later silent bug that ten years were actually for.

Collapse
 
infoinlet1 profile image
Info Inlet •

This is the sharpest addition anyone's made, because you turned the vibe into something checkable, which was exactly the gap I couldn't close in the post.

The durability-contract test is the one I'd tattoo on people. "Acked before saved can never pass CI twice" is the whole thing. The catch, and I say this having lived it, is that you can only write that test if you already had the judgment to know the failure exists. The test doesn't generate the insight, it preserves it. So it's less a substitute for the "no" and more a way to stop paying for the same "no" twice, which honestly might be the best you can do with judgment: catch it once by hand, then never by hand again.

The failure-story review is the part I want to push on teams. Reviewing the happy path is how the ack-before-persist bug sailed through in the first place, mine included. "What happens when the retry lands at exactly the wrong moment" should just be a required field on every PR. And your tell for it is perfect: if the answer needs a paragraph, that's the part you can't rent. I'm stealing that line, with credit.

"The three-weeks-later silent bug is what ten years were actually for" is a cleaner version of my own thesis than I wrote. Yeah. That's it.

Collapse
 
parsa4873fe3aa profile image
Parsa Mohammadi •

Agreed, and that means you only write the durability test for failures someone already recognized. The retry bug that burned a different team never makes it into a test. Neither does the one that hit this team two years ago, before anyone on it now was hired.

At Tomosu we work on the Production Reliability Index, which scores new changes against the incident history a team already has. So "never by hand again" also covers failures nobody remembers, including ones that never got a test. The first "no" still has to come from a person.

Thread Thread
 
contentclips_st profile image
ContentClips •

The blind spot on the far side of it: incidents that never got written up. Every team runs a shadow incident history β€” the retry bug someone fixed silently, the near-miss a senior caught by smell β€” and history-based scoring can only rank what got filed. The fault-injection fixture from earlier covers a different slice: it doesn't ask whether the failure is guarded, it asks whether anyone notices when it isn't. Random breakage plus one honest assertion tests the noticing itself. History tells the next team what to fear; the fixture tells you whether the skepticism still works. Both miss the silent fixes β€” which is why the first 'no' stays human.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

@parsa the production reliability index is the piece i didn't have a name for. my whole method is me walking the seams asking "what if this runs twice," but i can only ask that about failures i've personally been burned by. my skepticism is capped at my own scar tissue. scoring a change against the team's full incident history is how you borrow scars you never earned, including the ones that predate everyone currently on the team. that's the real gap in my approach and i'd use it. the "first no still has to come from a person" caveat is the important one though β€” the index can rank the fear, it can't originate it.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

@contentclips_st this is the sharper edge of it, yeah. history-based scoring can only ever see incidents that got filed, and in my experience the filed ones are the minority. the retry someone quietly fixed on a friday, the near-miss a senior killed by smell and never wrote up β€” those never make it into any history, so any tool built on that history is blind to them.

the distinction you drew is the part i'm keeping: history tells you what to fear, the fixture tests whether you'd even notice. i'd been mushing those into one thing. "does anything scream when it breaks" is a completely different question from "did we handle case X," and testing the noticing instead of the code is a framing i hadn't had. both still miss the silent fix, which is exactly why the first "no" stays human β€” a person smelling it is the only reason it ever got caught in the first place.

Collapse
 
contentclips_st profile image
ContentClips •

The 'stop paying for the same no twice' framing nails what I left fuzzy. One clause I'd add: expiry. A test encoding judgment is a claim about a system that drifts - when the retry path gets rewritten and everything still passes, that's not proof the reasoning holds, it's proof the scenario disappeared. So the failure-story field should do double duty: required on the PR, and written as a retirement note ('this no longer applies because X'). A durability contract nobody can retire turns into cargo cult faster than the judgment it froze.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

Expiry is the failure mode I hadn't named, and it's worse than the one I did, because a green vacuous test reads exactly like a green load-bearing one. The judgment didn't get verified, the scenario just quietly moved out and left the light on.

The retirement note doing double duty is the right shape. What I'd add: the note needs a trigger, not just a home. "This no longer applies because X" is only honest if something forces you to look when X changes β€” otherwise it rots at the same rate as the test it's guarding. The move I've landed on is making the test prove the scenario still exists before it proves the scenario is handled: an assertion that the retry path can still hit that window at all, so the day a rewrite makes it unreachable, the test goes red for "this can't happen anymore, come confirm that's true" instead of staying green and meaning nothing. Red-when-the-world-moved beats green-when-the-world-moved.

Which is really your whole point one level up: the drift isn't a test problem, it's the same judgment problem again, just deferred. Someone still has to notice the terrain changed. All the contract does is decide whether that noticing happens on a schedule or during the outage. "Cargo cult faster than the judgment it froze" is the exact risk β€” a frozen no that nobody's allowed to thaw is just superstition with a green check.

Thread Thread
 
contentclips_st profile image
ContentClips •

'Red-when-the-world-moved' is the right failure direction - and that existential assertion is really one line of fault injection distilled, which makes it the cheapest chaos discipline a team can adopt: no platform, just a fixture that pins the window open in CI. On the scheduling question, I'd push the ownership: the person who wrote the judgment is structurally biased toward keeping it - sunk 'no'. Whoever touches the module next should own the confirm-vs-retire call; note-plus-trigger turns it into a two-minute decision, which shifts that noticing from 'during the outage' to 'on a schedule' without adding a meeting. The contract then does the whole job you gave it: schedule the noticing AND hand it to the person with the least sunk cost in the old terrain.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

Handing the retire call to whoever touches the module next, not whoever wrote the "no," is what makes it self-cleaning. The author's the worst owner of that decision, they'll defend a scenario that moved out three rewrites ago because it cost them an outage to learn it. Fresh eyes actually read the note and ask if it's still true.

And that's the loop closed. Every layer we added just moved the judgment around. This is the first one that doesn't move it, it hands it to the person with the least reason to lie about it.

Best thread I've had on here, honestly. Turned a confession into something a team can run Monday.

Collapse
 
slabb profile image
Sam LABBE •

The origin story you picked is doing more work than you're claiming for it. "The 'got it' went out, the save never happened, and there was nothing in the logs" β€” that wasn't only a judgment miss, it was a record miss. Two events that should have been reconcilable β€” what you promised the customer vs. what actually persisted β€” and nobody could replay them afterward. Judgment is how you catch it the second time. The record is how it gets caught at all.

So I'd add a quiet fourth member to author / skeptic / human: the record. A skeptic's "no" that isn't written down is indistinguishable from a "no" that never happened β€” and a human who "owns the call" with no trace of the call owns it rhetorically. The cheap version is a twist on your own ritual: you already say "here's how this loses money" out loud, wrong most of the time. Write the sentence down, timestamped, next to the diff it objected to. Do that for a year and "wrong most of the time" stops being a shrug and becomes calibration β€” twelve real disasters out of forty objections is a track record, and it's the only line on the rewritten CV nobody can argue with.

Full disclosure: this is the pattern behind the flight-recorder journal I build, so consider the source. But I think you've underpriced your own skill. Judgment you can't show is indistinguishable from luck. The no is the skill β€” the receipt is what makes it one.

To your question, then: the one I'd keep is deciding what deserves a receipt.

Collapse
 
infoinlet1 profile image
Info Inlet •

ok this is the best pushback i've gotten and i think you're right. i was treating it as a pure judgment miss and it wasn't. the judgment failure and the record failure were two different failures stacked on top of each other, and honestly the record one was the worse of the two. if the promise and the persist had been reconcilable afterward, someone finds that customer in a day. without it they just vanish, and the lesson vanishes with them. i felt the second failure and filed it under the first. that's on me.

the "no that isn't written down is indistinguishable from a no that never happened" line is going to live in my head for a while. because yeah, i've been quietly proud of catches i can't actually prove i made. that's not a track record, that's a story i tell myself. twelve out of forty is a number. a vibe isn't.

and i'll take the hit on underpricing it. "judgment you can't show is indistinguishable from luck" is exactly the hole in my own post. i said the only score that's mine is the disasters i saw coming, then gave no way to count them. the receipt is what turns it from a feeling into a score.

the self-plug is fine, you flagged it and it earned its place, so no complaints there.

the one spot i'd push back a little: i don't think the record replaces the no, and i don't think you're saying it does, but worth keeping straight. a receipt on a bad call is still a bad call, just a legible one. writing it down doesn't make the judgment good, it makes it checkable, which is the thing that lets it get good over time. so maybe the honest shape is the no is the skill, the receipt is what lets you find out if the skill is real.

"deciding what deserves a receipt" is a great answer because it's recursive, it's a judgment call about judgment. you can't write everything down or the record's just noise. deciding what's worth the ink is the same muscle one level up. stealing that framing, with attribution.

Collapse
 
slabb profile image
Sam LABBE •

The line that stayed with me is "nothing in the logs." Judgment catches what green tests can't β€” agreed. But the facility that answers "what actually happened to that customer" is separate from judgment, and it's buildable: the ack-before-persist decision is an event like any other, and it can be recorded in a form that survives the post-mortem. Judgment finds the bug; the record decides whether the next engineer believes you.

Thread Thread
 
infoinlet1 profile image
Info Inlet •

Yeah, I'm fully bought in now. The thing that finally clicked reading this: judgment is a property of a person, and the record is what lets it outlive that person being in the room. "Nothing in the logs" didn't just hide the bug from me that night β€” it meant nobody who came after could reconstruct the call either. The ack-before-persist decision happened, it was a real event with a real cause, and the system kept no memory that it had ever been made. The bug was recoverable in principle and unrecoverable in practice, and that gap is entirely a record problem, not a judgment one.

"The record decides whether the next engineer believes you" is the sharpest way anyone's put it in this whole thread. Because that's the failure mode of pure judgment β€” it's non-transferable. I can't hand my hunch to the junior, and I can't prove it to the skeptic, and six months later I can't even prove it to myself. An un-recorded no dies with the shift it happened on. The receipt is what lets a decision survive turnover, survive the post-mortem, survive me forgetting I ever made it.

So where I've landed: author produces, skeptic objects, human owns the call β€” and the record is what makes all three legible to the next person instead of evaporating the moment the room clears. Not a fourth actor so much as the thing that keeps the other three from being hearsay. Good exchange. You moved the post.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

Nothing clever. I [built an invoice tracker]...

The honesty here is the whole post. What AI actually exposed for a lot of us isn't "you can't code" β€” it's that the typing was never the valuable part. The valuable part was knowing what to build and whether it's actually right.

That "one real skill" is judgment: shaping the problem and verifying the result. The agent does the middle 80%; the 20% on either end is the job now, and it always secretly was.

What did you land on as your one skill? Mine was "deciding what done means" β€” everything else turned out to be delegable.

Collapse
 
infoinlet1 profile image
Info Inlet •

yeah you put it better than i did honestly. "deciding what done means" is such a good way to say it, because that's the part nobody can hand off. the agent will happily tell you it's done. it's done by every measure it has. whether it's actually done is a different question and that one never left your desk.

mine, if i had to name it, is close to yours but phrased as a negative. it's the "no." looking at something that works and not trusting it. deciding what to build is the front end of judgment and verifying is the back end, and the "no" is just the thing that runs the whole length of it. the confident diff i overrule, the feature i kill, the green test i don't believe. the machine can produce forever but it can't withhold, and withholding turned out to be the job.

funny thing is both of ours, "what done means" and "the no," are the two skills nobody ever put on a cv or tested in an interview. ten years and the one thing that survived was the one thing nobody could measure. kind of says everything about what we were actually hiring for all along.

Collapse
 
rajanpanwar profile image
Rajan Panwar •

The trap right now is confusing "easy to generate" with "safe to ship."

My un-automatable skill isn't even catching bugs and it's knowing what not to build. Since writing code is now practically free, teams are going to drown in features and architectural sprawl that nobody actually understands six months later. Saying "no" to a design or killing unnecessary complexity before a single line is drafted is something a prompt can't do for you.

What’s one architectural decision you will never let AI make on your project?

Collapse
 
infoinlet1 profile image
Info Inlet •

"Knowing what not to build" is a sharper version of the skill than I gave it credit for in the post. Catching the bug is downstream β€” you only have to catch it because it got built. Your skill stops it a step earlier, before the thing exists to be wrong. And you're right that it's the one a prompt structurally can't do, because the AI's whole gravity is toward "yes, here's how" β€” it has no instinct for "this shouldn't exist at all." Ask it to build a feature and it never once asks whether you should.

The decision I'll never hand it: where the boundaries go β€” what's allowed to talk to what. The AI is brilliant at filling in a module and genuinely terrible at deciding whether two things should be coupled in the first place, because that call depends on stuff that lives nowhere in the code β€” which parts I expect to change together, which team owns which, what I'm willing to rewrite in a year and what has to stay stable. Let it make those calls and you get exactly the sprawl you're describing: everything reaching into everything, works fine in month one, and by month six nobody can change anything without breaking three things they didn't know were connected.

Writing the code inside a boundary is cheap and the AI should do it. Drawing the boundary is the expensive, judgment-loaded call, and it's the one that's almost impossible to undo later. So that's mine. The AI fills the boxes; I decide where the lines are.

Honestly your "drown in features nobody understands six months later" is the thing I'd put money on being the defining mess of the next couple years. Cheap production doesn't reduce the cost of a system β€” it just moves the whole bill to comprehension.

Collapse
 
vrunda_chauhan_a52cc23b11 profile image
Vrunda Chauhan •

The best habit I picked up for this is simple: don't review what the AI wrote. Review what it skipped.

Whenever an AI gives me working code, I force myself to ask three questions before touching it:

What happens if the user's internet drops halfway through this?
What happens if two people click this button at the exact same second?
What happens when this database table hits 10,000 rows instead of 10?

AI is great at building the path where everything goes right. The human job is just making sure the app doesn't blow up when things go wrong.

Collapse
 
infoinlet1 profile image
Info Inlet •

"Review what it skipped" might be the best four-word version of the whole post. Because that's the asymmetry exactly β€” the AI shows you the path it built, confidently and completely, and the danger is never in what's on the screen. It's in what isn't. And nothing in the output points at the hole, because from the model's side the hole doesn't exist. It answered the question you asked; the bug is in the question you didn't.

Your three questions are great, and I'd notice what they have in common: not one of them is answerable by reading the code. Internet dropping halfway, two clicks in the same second, 10 rows becoming 10,000 β€” those all live in the world the code runs in, not the code itself. That's why the AI misses them and why static review misses them too. You're not checking the picture, you're checking the frame around it.

The one I'd add to your list, since it's the one that got me: "what happens if this exact request arrives twice?" Retries, double-submits, a webhook that fires again β€” same family as your two-clicks question, and it's quietly behind half the money-losing bugs I've seen. The happy path handles it once, beautifully, and falls apart the second time.

The only thing I'd gently push on is "the human job is just making sure it doesn't blow up when things go wrong" β€” I'd drop the "just." That's not the small remaining chore after the AI did the real work. That is the work now. The AI took the part that was always easy and left us the part that was always the actual job. Your three questions aren't a safety check on top of development. They're the development.

Collapse
 
kartik-nvjk profile image
Kartik N V J K •

The "is this going to quietly lose someone money in three weeks" line is the whole thing for me. Production is easy to check, it either runs or it does not, but that slow judgment call almost never shows up in a test, which is why a first-try-working invoice tracker can still be wrong in a way nobody sees for a month. How do you try to surface that kind of delayed failure before it ships?

Collapse
 
infoinlet1 profile image
Info Inlet •

honestly that line still makes me nervous and you nailed why. green tests will wave that bug through all day because it's just not a test-shaped problem.

so i don't really test for it. the thing that's helped most is boring: i stopped looking inside the function and started looking at the edges. the three week bug almost never lives in the logic, it lives where my code hands off to something else. a network call, a queue, a db write, a client that might retry. so before i trust anything i just walk those handoffs and ask the same dumb question, what if this step runs twice, or what if it runs and the next one doesn't. that ack-before-persist bug was exactly that. the code was fine, the order across the boundary wasn't.

the other habit is i make myself actually say the sentence out loud, "here's how this loses someone money." if i can't finish it i don't really understand the change yet, i just know it runs. most of the time the sentence is nonsense and there's nothing there, fine. the one time it writes itself easily is the one you go fix.

and i've basically given up reviewing my own stuff for this. whoever wrote it is the worst person to catch it because they already decided it's right, that's why it came out that way. so i let someone (or something) else just try to break the happy path. replay the request, send it twice, kill it halfway through. not "does it work" but "show me where it half worked."

none of it catches everything though. the whole definition of these bugs is they slip whatever net you knew to build. i've just stopped letting "all tests pass" stand in for "someone actually tried to make this thing lose money." not the same sign off.

Collapse
 
nullandvoid_ profile image
Jyanthi •

The distinction between production and judgment really hits home. AI can generate increasingly impressive solutions, but knowing when not to trust a solution is a different skill entirely. The β€œauthor, skeptic, human” model is a fascinating way to think about where engineering is heading.

Collapse
 
infoinlet1 profile image
Info Inlet •

Thanks Jyanthi. The thing that gets me is the two skills feel identical from the inside β€” a confident wrong answer and a confident right one arrive in the exact same voice. That's why the "don't trust it" can't live in the same place that generates the answer. It has to be a separate seat, or it just agrees with itself.