DEV Community

Cover image for I let AI write my code for a month. The junior devs caught what it broke — I didn't.
Rudratosh Shastri
Rudratosh Shastri

Posted on

I let AI write my code for a month. The junior devs caught what it broke — I didn't.

I'll say the unpopular part first: for one month I leaned on an AI coding agent for almost everything, and the bugs it introduced were caught by the two most junior people on the team. Not by me. Not by the tests. By the two people the internet keeps telling us AI is about to replace.

I want to talk about why, because the "AI made me 10x faster" posts are all true and all incomplete. The speed is real. The failure mode is also real, and it's more dangerous than the one everyone warns you about.

The bug you expect vs. the bug you get

Everyone braces for the obvious failure: the AI writes something that doesn't compile, or hallucinates a function that doesn't exist. That bug is fine. It's loud. The compiler screams, the test goes red, you fix it in ten seconds. Loud bugs are cheap.

The bugs that got me were the opposite. They compiled. They passed the tests I had. They looked exactly like code I would have written — because they were trained on code I would have written. They were plausible. And plausible is the single most expensive property a bug can have, because plausibility is what switches your review brain off.

Three examples from that month:

  • A caching helper that was correct except it keyed the cache on a mutable object, so under concurrency it occasionally served one user another user's result. Passed every test. There was no concurrency in the tests.
  • A retry wrapper that retried on all exceptions, including a validation error that would never succeed — turning a clean 400 into a 30-second hang and three log lines that looked like a network blip.
  • A refactor that "simplified" a permission check by collapsing two conditions into one, which read beautifully and quietly widened access by one role.

Every one of those is the kind of thing I'd catch instantly in a stranger's PR. In AI output, I skimmed right past them. Three times.

Why I missed them and they didn't

Here's the honest mechanism, and it's not about skill.

When I write code, I've already argued with myself about the edge cases on the way there. My review of my own code is a re-run of an argument I remember having. When the AI writes it, that argument never happened. There's no memory of the reasoning to re-run — just fluent output that looks like the conclusion of reasoning. So I reviewed the look of it, matched it against "is this how I'd write it," got a yes, and moved on. Fast, confident, wrong.

The juniors did the opposite, for the exact reason they're junior: they don't trust code they don't understand yet. They read the caching helper line by line because they had to, to learn it. And reading line by line is precisely what catches a mutable cache key. Their inexperience forced the slow path. My experience let me take the fast one. The fast one is where plausible bugs live.

That inverts the whole "juniors are obsolete now" narrative. In an AI-heavy workflow, the person who reads every line because they can't yet skim is doing the most valuable job on the team. The senior who trusts their pattern-match is the liability.

What I actually changed

I didn't stop using AI. It genuinely is faster, and for the loud-bug category — boilerplate, glue, one-shot scripts — it's close to free. What I changed was the review contract:

  1. AI code gets stricter review than human code, not looser. The instinct is the reverse ("it's probably fine, the model is good now"). Invert it. No argument happened in its head; the argument has to happen in yours.
  2. Review it like a stranger wrote it. Because a stranger did. Strip the "this looks like me" reflex — that reflex is the exploit.
  3. The tests it passes are the tests you already had. AI is very good at satisfying the existing suite and silent about the case the suite never covered. Ask, every time: what input is not in these tests? That's where the plausible bug is hiding.
  4. Let the person reading slowly win the argument. When a junior says "wait, why does this retry on everything?" — that's not them being slow. That's the review working. Protect that.

The part I keep coming back to

We keep framing AI-and-juniors as a replacement question: does the model do the junior's job now? A month of my own screwups says the framing is wrong. The model does the typing. It does not do the doubting. And doubting — reading the line you'd normally skip, distrusting the code that looks right — turns out to be the job that actually protects production.

The junior who slows down to understand isn't behind the curve. In a world where the code writes itself fluently and confidently and sometimes wrongly, they're the curve.

AI didn't make my juniors obsolete this month. It made them essential, and it made me the weak link. I'd rather say that out loud than post another 10x-speed screenshot.


Be honest: has AI-written code ever slipped a bug past you that you'd have caught in a human's PR? What was it, and who found it? 👇

I write about building with AI and the honest ways it breaks. Follow me here if that's your lane. 👋

Top comments (32)

Collapse
 
vladzoff profile image
Vlad Zoff •

The point about tests only covering the questions you already wrote down has another version at the architecture level. I've seen agent changes keep every test green while quietly bypassing an existing module boundary or adding a dependency in the wrong direction. That's harder to notice because the behavior is still correct.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

agent changes keep every test green while quietly bypassing an existing module boundary or adding a dependency in the wrong direction

This is the scarier sibling of the point, and you've named it precisely: the behavior is correct, so no test fires, but the architecture quietly rots. Tests assert behavior; they almost never assert structure.

The only thing that's caught this for me is making structure itself testable — dependency-direction lint / import boundaries (dependency-cruiser, an ArchUnit-style check) — so "UI imported from infra" fails CI even when the feature works. An agent can't argue with a boundary encoded as a rule it can see.

Have you found a lint/arch check that catches the "right direction, wrong dependency" case, or is it still human review doing the catching?

Collapse
 
vladzoff profile image
Vlad Zoff •

Yes, dependency-direction checks are a good first layer. The harder case is exactly the one you mentioned: the dependency follows the allowed direction, but the specific dependency is still wrong for that boundary.
For example, route -> service may be valid, while one particular route should not depend on a specific service, or a service may be allowed to use a repository but not bypass another application boundary. That's difficult to express as a generic import-direction rule because the answer depends on the project's actual architecture.
That's the part I'm working on (with Guard): discover the existing structure first, then let the project define those more specific boundaries as rules/contracts. Human review still makes sense for cases where the question is really about intent rather than something the repository can express deterministically.

Thread Thread
 
rudratosh profile image
Rudratosh Shastri •

the dependency follows the allowed direction, but the specific dependency is still wrong for that boundary

That's the gap generic lint can't close. Direction rules encode the layering, not the architecture. "Routes may call services" says nothing about this route calling that service.

Discovering the existing structure first and then letting the project declare contracts sounds like the right order, because:

  • Hand-written rules from scratch never get written
  • Inferred rules with no human sign-off will happily bless the existing mistakes
  • Inferred, then confirmed gives you a baseline someone actually agreed to

That second point is my real question. When Guard discovers the current structure, how does it tell an intentional boundary from an accident that just happens to be consistent so far?

Collapse
 
elijahbrown profile image
Elijah Brown •

The plausible-bug point lands. One case I'd add to 'what input is not in these tests' for AI-written forms: a well-formed email on a domain with no MX, and a phone typed in local format for the wrong country. Format checks pass both; a server-side parse catches them.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

Exactly — those are great examples of the “looks valid, behaves wrong” class of bugs I was getting at.

A syntactically valid email isn't necessarily deliverable, and a correctly formatted phone number isn't necessarily valid for the user's country. Those are precisely the cases that a happy-path test suite can miss because the input satisfies the validator while violating the real-world constraint.

And the server-side validation point is important too: the client can check shape, but the backend needs to validate the actual semantics before trusting the value. That's another good example of why “the tests pass” isn't the same as “the input space is covered.”

Collapse
 
elijahbrown profile image
Elijah Brown •

Agreed, and both make cheap test fixtures: a domain that publishes a Null MX, and a national-format number sent with no country, since the same digits can be valid in one country and invalid in another.

Thread Thread
 
rudratosh profile image
Rudratosh Shastri •

Exactly. Those are great examples because neither requires an exotic edge case — they’re cheap, deterministic fixtures that expose whether the implementation actually understands the semantics.

That’s the bigger lesson with AI-generated code: a handful of deliberately adversarial test cases can reveal gaps that a large pile of happy-path tests completely misses.

Thread Thread
 
elijahbrown profile image
Elijah Brown •

Exactly. Small deterministic cases like those are a good way to expose semantic gaps that happy-path tests miss.

Thread Thread
 
rudratosh profile image
Rudratosh Shastri •

Small deterministic cases … expose semantic gaps that happy-path tests miss

Exactly — and the deterministic part is what makes them cheap to keep: a Null-MX domain and a national-format number with no country code never go stale, never flake, and each encodes a real-world constraint the validator doesn't. I've started keeping a little "adversarial fixtures" file that's just these — the cases I'd never have written if I hadn't been burned. What's the one you reach for most?

Thread Thread
 
elijahbrown profile image
Elijah Brown •

The one I reach for most is a reserved phone fixture that is syntactically valid but intentionally non-routable, paired with a domain that publishes Null MX; it catches normalization and policy mistakes without touching live data.

Thread Thread
 
elijahbrown profile image
Elijah Brown •

The one I reach for most is a replay of an old request against a changed schema or contract, with assertions on both the result and the error state. It catches the quiet compatibility break that a fresh happy-path fixture tends to miss.

Thread Thread
 
rudratosh profile image
Rudratosh Shastri •

a replay of an old request against a changed schema or contract, with assertions on both the result and the error state

That's the one most suites miss, and it's the most dangerous gap because it fails silently — the happy path still passes while the contract quietly moved underneath it. Asserting the error state, not just the result, is the part people skip and exactly the part that catches it.

Between this and the Null-MX / non-routable-phone pair, there's a theme: the fixtures that earn their keep all encode a real-world constraint the validator doesn't know about — deliverability, routability, backward compatibility. Cheap, deterministic, never flaky. That's why they belong in a dedicated adversarial-fixtures file: they're precisely the cases AI-written code (and I) tend to assume away. Adding the contract-replay one — thanks.

Thread Thread
 
elijahbrown profile image
Elijah Brown •

That contract-drift fixture is the one I’d keep closest to the request corpus: it catches a change that a happy-path test can hide. Asserting the expected error shape as well as the success shape makes the failure visible.

Thread Thread
 
rudratosh profile image
Rudratosh Shastri •

Asserting the expected error shape as well as the success shape makes the failure visible.

That's the whole move. A happy-path test asserts the result; a contract-drift test asserts the error, and that's the half that catches a silent break. Keeping it next to the request corpus (not off in a separate "edge cases" file) is smart too — it travels with the thing it protects. Between this, the Null-MX domain and the non-routable phone, the pattern's clear: the fixtures worth keeping encode a constraint the validator doesn't know about. Stealing the error-shape assertion.

Thread Thread
 
elijahbrown profile image
Elijah Brown •

Exactly. Keeping the fixture beside the request corpus makes that constraint visible when the contract changes.

Thread Thread
 
elijahbrown profile image
Elijah Brown •

That is a good home for it. Keeping the replay beside the other deterministic fixtures should make contract drift visible before it reaches the UI.

Thread Thread
 
elijahbrown profile image
Elijah Brown •

That is the part that makes the fixture worth keeping. A replay beside the request corpus keeps the changed contract visible, while the error assertion stops a silent success from looking healthy.

Thread Thread
 
rudratosh profile image
Rudratosh Shastri •

Perfectly put — that's the resting place for the whole thread: a replay beside the request corpus keeps the drift visible, and the error-shape assertion stops a silent success from passing as healthy. Small, deterministic, lives next to the thing it protects. Genuinely sharpened how I'll set these up — thanks for the back-and-forth.

Thread Thread
 
elijahbrown profile image
Elijah Brown •

Thanks, that is exactly the kind of small fixture I want beside the request corpus. It keeps contract changes visible without adding much maintenance.

Collapse
 
respect17 profile image
Kudzai Murimi •

"My review of my own code is a re-run of an argument I remember having" explains exactly why skimming AI output is so dangerous, there's no argument to re-run because it never happened in your head. The mutable cache key bug is a perfect example, that's the kind of thing that only shows up under load, long after the PR was approved.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

Exactly. That’s the uncomfortable part: AI gives you the appearance of reasoning without you having experienced the reasoning.

The mutable cache key was a good example because nothing looks obviously wrong in the code. The missing question is what happens under conditions the test suite never created — concurrency, unusual inputs, permission boundaries, failure paths. That’s where “looks right” becomes dangerous.

Collapse
 
danielecangi profile image
DaC •

Interesting experiment, but before drawing conclusions I'd be curious about the actual setup. Which coding agent were you using, and how? copilot agent mode, claude code, Codex, something else? Was it mainly generating code from prompts, or could it inspect the repo, run and add tests exc exc? I ask because those are very different workflows. a month is enough to expose failure modes, but probably not enough to generalize about “AI coding agents” as a category. It took me years of working with these models

Collapse
 
rudratosh profile image
Rudratosh Shastri •

before drawing conclusions I'd be curious about the actual setup. Which coding agent were you using, and how?

Fair question, and you're right that it changes everything — so let me be concrete instead of hand-wavy.

Setup was the capable end, not prompt-in-a-box: repo-aware agents (Claude Code and Codex, some Cursor) that could read the whole codebase, run the test suite, and add their own tests. So this isn't the weak-setup failure mode — if anything it's the opposite, which matters for the conclusion.

And your bigger point is the one I'd defend hardest, by conceding it:

a month is enough to expose failure modes, but probably not enough to generalize about "AI coding agents" as a category

Agreed — n=1 dev, one month, a handful of projects is a failure-mode report, not a category verdict. I'm not claiming "agents write bad code." The narrower claim, and I think a setup-independent one: the more repo-aware and capable the agent, the more plausible its wrong code — and the faster my own review discipline atrophied to match it. The good setup made this worse, not better, because the output was convincing enough that "looks right" quietly replaced "I actually reasoned through this."

That's why the junior devs caught what I didn't: they hadn't outsourced the argument, so they still ran it.

You said it took you years with these models — genuinely curious: what changed in how you review agent output over that time? Did you land on a discipline that survives the plausibility, or was it mostly that you stopped trusting "looks right" the hard way?

Collapse
 
maximin_mxn_6ce1a3054be6d profile image
Maximin •

The most useful part is the review mechanism: AI removes the memory of the argument that should have happened before the code existed. Treating generated code like a stranger’s PR, then testing the missing inputs deliberately, is a practical way to recover that skepticism.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

Exactly. I think the biggest hidden cost of AI-generated code is not the code itself — it's the missing conversation that normally happens before a human writes it.

When we write code ourselves, we naturally carry context: why this approach, what assumptions were made, what edge cases were ignored. AI skips that history, so reviewing it like an unfamiliar PR forces those questions back into the process.

The "stranger's PR" mindset has become one of my favorite ways to use AI safely: don't ask "does this work?" first, ask "what would I challenge if someone else opened this PR?"

Please Follow ♥️ if you like my post!

Collapse
 
kyisaiah47 profile image
kyisaiah47 •

The retry example leaves me wondering whether the validation error had a distinct type or status that the wrapper could inspect before retrying.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

Great catch. That example was intentionally simplified, but you are right — a real retry wrapper should absolutely distinguish between error categories before deciding to retry.

A validation error (bad input, schema mismatch, missing required field) usually needs correction, not another attempt. Retrying those just burns tokens and can even hide the real issue.

The safer pattern is closer to:

  • retry transient failures (timeouts, rate limits, temporary upstream errors)
  • stop and surface deterministic failures (validation, permission, business rules)
  • log the context that led to the failure

The interesting part with AI agents is that this classification layer becomes even more important because the agent can generate the retry logic itself — and that logic also needs review.

Please Follow ♥️ if you like my post!

Collapse
 
kozmonot20 profile image
kozmonot20 •

The Day My Python Script Almost Sent My PC to Mars

ai

python

opensource

monitoring
🚀 The Day My Python Script Tried to Burn Down the House (A Qwen & Local AI Story)
So, there I was, a 12-year-old developer trying to optimize my emektar PC (i5-4200U, 4GB RAM) because local AI models like Qwen2:0.5b are amazing but they eat RAM for breakfast. 🧠💻
I decided to write a super-advanced "All-in-One Python Booster". The plan was simple:

Use powercfg to force Windows into Extreme Performance Mode.
Use psutil to aggressively terminate any background apps sucking up my precious 4GB RAM. I opened my terminal, fired up the script, and whispered to myself: "Let's fly." 🚀 Mistake Number 1: I forgot to create a whitelist for critical Windows services. Mistake Number 2: I left my experimental rocket science chemicals (potassium nitrate) on the desk next to the laptop fan. 🧪💨 Within three seconds, my Python loop went completely rogue. It detected Windows' own kernel processes as "background garbage" and started terminating them. The screen went crazy: • Blue Screen 1: IRQL_NOT_LESS_OR_EQUAL (Windows was panicking) 🛑 • Blue Screen 2: CRITICAL_PROCESS_DIED (The kernel was literally dying) 💀 • Blue Screen 3: DPC_WATCHDOG_VIOLATION (The watchdog dog bit the system) 🐕 As the CPU hit 1000% load trying to execute my aggressive loops, the laptop fan started spinning so fast it created a localized vortex. It literally sucked the potassium nitrate dust straight from the desk into the exhaust vents. 🌪️ Suddenly, my PC decided to stop being a computer and started practicing to become a SpaceX rocket. A literal flash of fire shot out of the side vents. My laptop "osurarak" breathed its last breath of smoke and fell into the nearby barbecue grill. Burned, fried, and completely dead. 😭🔥💨 The Moral of the Story: Windows might have died, and my hardware might be ashes, but thanks to the cyber-gods, my code was already pushed to GitHub! 🦾😎 Always backup your code, whitelists are important, and please, keep your rocket fuel away from your while True loops. 73 to all amateur radio and Python enthusiasts out there! 📡✈️

Some comments may only be visible to logged-in visitors. Sign in to view all comments.