I'll say the unpopular part first: for one month I leaned on an AI coding agent for almost everything, and the bugs it introduced were caught by the two most junior people on the team. Not by me. Not by the tests. By the two people the internet keeps telling us AI is about to replace.
I want to talk about why, because the "AI made me 10x faster" posts are all true and all incomplete. The speed is real. The failure mode is also real, and it's more dangerous than the one everyone warns you about.
The bug you expect vs. the bug you get
Everyone braces for the obvious failure: the AI writes something that doesn't compile, or hallucinates a function that doesn't exist. That bug is fine. It's loud. The compiler screams, the test goes red, you fix it in ten seconds. Loud bugs are cheap.
The bugs that got me were the opposite. They compiled. They passed the tests I had. They looked exactly like code I would have written — because they were trained on code I would have written. They were plausible. And plausible is the single most expensive property a bug can have, because plausibility is what switches your review brain off.
Three examples from that month:
- A caching helper that was correct except it keyed the cache on a mutable object, so under concurrency it occasionally served one user another user's result. Passed every test. There was no concurrency in the tests.
- A retry wrapper that retried on all exceptions, including a validation error that would never succeed — turning a clean 400 into a 30-second hang and three log lines that looked like a network blip.
- A refactor that "simplified" a permission check by collapsing two conditions into one, which read beautifully and quietly widened access by one role.
Every one of those is the kind of thing I'd catch instantly in a stranger's PR. In AI output, I skimmed right past them. Three times.
Why I missed them and they didn't
Here's the honest mechanism, and it's not about skill.
When I write code, I've already argued with myself about the edge cases on the way there. My review of my own code is a re-run of an argument I remember having. When the AI writes it, that argument never happened. There's no memory of the reasoning to re-run — just fluent output that looks like the conclusion of reasoning. So I reviewed the look of it, matched it against "is this how I'd write it," got a yes, and moved on. Fast, confident, wrong.
The juniors did the opposite, for the exact reason they're junior: they don't trust code they don't understand yet. They read the caching helper line by line because they had to, to learn it. And reading line by line is precisely what catches a mutable cache key. Their inexperience forced the slow path. My experience let me take the fast one. The fast one is where plausible bugs live.
That inverts the whole "juniors are obsolete now" narrative. In an AI-heavy workflow, the person who reads every line because they can't yet skim is doing the most valuable job on the team. The senior who trusts their pattern-match is the liability.
What I actually changed
I didn't stop using AI. It genuinely is faster, and for the loud-bug category — boilerplate, glue, one-shot scripts — it's close to free. What I changed was the review contract:
- AI code gets stricter review than human code, not looser. The instinct is the reverse ("it's probably fine, the model is good now"). Invert it. No argument happened in its head; the argument has to happen in yours.
- Review it like a stranger wrote it. Because a stranger did. Strip the "this looks like me" reflex — that reflex is the exploit.
- The tests it passes are the tests you already had. AI is very good at satisfying the existing suite and silent about the case the suite never covered. Ask, every time: what input is not in these tests? That's where the plausible bug is hiding.
- Let the person reading slowly win the argument. When a junior says "wait, why does this retry on everything?" — that's not them being slow. That's the review working. Protect that.
The part I keep coming back to
We keep framing AI-and-juniors as a replacement question: does the model do the junior's job now? A month of my own screwups says the framing is wrong. The model does the typing. It does not do the doubting. And doubting — reading the line you'd normally skip, distrusting the code that looks right — turns out to be the job that actually protects production.
The junior who slows down to understand isn't behind the curve. In a world where the code writes itself fluently and confidently and sometimes wrongly, they're the curve.
AI didn't make my juniors obsolete this month. It made them essential, and it made me the weak link. I'd rather say that out loud than post another 10x-speed screenshot.
Be honest: has AI-written code ever slipped a bug past you that you'd have caught in a human's PR? What was it, and who found it? 👇
I write about building with AI and the honest ways it breaks. Follow me here if that's your lane. 👋
Top comments (32)
The point about tests only covering the questions you already wrote down has another version at the architecture level. I've seen agent changes keep every test green while quietly bypassing an existing module boundary or adding a dependency in the wrong direction. That's harder to notice because the behavior is still correct.
This is the scarier sibling of the point, and you've named it precisely: the behavior is correct, so no test fires, but the architecture quietly rots. Tests assert behavior; they almost never assert structure.
The only thing that's caught this for me is making structure itself testable — dependency-direction lint / import boundaries (dependency-cruiser, an ArchUnit-style check) — so "UI imported from infra" fails CI even when the feature works. An agent can't argue with a boundary encoded as a rule it can see.
Have you found a lint/arch check that catches the "right direction, wrong dependency" case, or is it still human review doing the catching?
Yes, dependency-direction checks are a good first layer. The harder case is exactly the one you mentioned: the dependency follows the allowed direction, but the specific dependency is still wrong for that boundary.
For example,
route -> servicemay be valid, while one particular route should not depend on a specific service, or a service may be allowed to use a repository but not bypass another application boundary. That's difficult to express as a generic import-direction rule because the answer depends on the project's actual architecture.That's the part I'm working on (with Guard): discover the existing structure first, then let the project define those more specific boundaries as rules/contracts. Human review still makes sense for cases where the question is really about intent rather than something the repository can express deterministically.
That's the gap generic lint can't close. Direction rules encode the layering, not the architecture. "Routes may call services" says nothing about this route calling that service.
Discovering the existing structure first and then letting the project declare contracts sounds like the right order, because:
That second point is my real question. When Guard discovers the current structure, how does it tell an intentional boundary from an accident that just happens to be consistent so far?
The plausible-bug point lands. One case I'd add to 'what input is not in these tests' for AI-written forms: a well-formed email on a domain with no MX, and a phone typed in local format for the wrong country. Format checks pass both; a server-side parse catches them.
Exactly — those are great examples of the “looks valid, behaves wrong” class of bugs I was getting at.
A syntactically valid email isn't necessarily deliverable, and a correctly formatted phone number isn't necessarily valid for the user's country. Those are precisely the cases that a happy-path test suite can miss because the input satisfies the validator while violating the real-world constraint.
And the server-side validation point is important too: the client can check shape, but the backend needs to validate the actual semantics before trusting the value. That's another good example of why “the tests pass” isn't the same as “the input space is covered.”
Agreed, and both make cheap test fixtures: a domain that publishes a Null MX, and a national-format number sent with no country, since the same digits can be valid in one country and invalid in another.
Exactly. Those are great examples because neither requires an exotic edge case — they’re cheap, deterministic fixtures that expose whether the implementation actually understands the semantics.
That’s the bigger lesson with AI-generated code: a handful of deliberately adversarial test cases can reveal gaps that a large pile of happy-path tests completely misses.
Exactly. Small deterministic cases like those are a good way to expose semantic gaps that happy-path tests miss.
Exactly — and the deterministic part is what makes them cheap to keep: a Null-MX domain and a national-format number with no country code never go stale, never flake, and each encodes a real-world constraint the validator doesn't. I've started keeping a little "adversarial fixtures" file that's just these — the cases I'd never have written if I hadn't been burned. What's the one you reach for most?
The one I reach for most is a reserved phone fixture that is syntactically valid but intentionally non-routable, paired with a domain that publishes Null MX; it catches normalization and policy mistakes without touching live data.
The one I reach for most is a replay of an old request against a changed schema or contract, with assertions on both the result and the error state. It catches the quiet compatibility break that a fresh happy-path fixture tends to miss.
That's the one most suites miss, and it's the most dangerous gap because it fails silently — the happy path still passes while the contract quietly moved underneath it. Asserting the error state, not just the result, is the part people skip and exactly the part that catches it.
Between this and the Null-MX / non-routable-phone pair, there's a theme: the fixtures that earn their keep all encode a real-world constraint the validator doesn't know about — deliverability, routability, backward compatibility. Cheap, deterministic, never flaky. That's why they belong in a dedicated
adversarial-fixturesfile: they're precisely the cases AI-written code (and I) tend to assume away. Adding the contract-replay one — thanks.That contract-drift fixture is the one I’d keep closest to the request corpus: it catches a change that a happy-path test can hide. Asserting the expected error shape as well as the success shape makes the failure visible.
That's the whole move. A happy-path test asserts the result; a contract-drift test asserts the error, and that's the half that catches a silent break. Keeping it next to the request corpus (not off in a separate "edge cases" file) is smart too — it travels with the thing it protects. Between this, the Null-MX domain and the non-routable phone, the pattern's clear: the fixtures worth keeping encode a constraint the validator doesn't know about. Stealing the error-shape assertion.
Exactly. Keeping the fixture beside the request corpus makes that constraint visible when the contract changes.
That is a good home for it. Keeping the replay beside the other deterministic fixtures should make contract drift visible before it reaches the UI.
That is the part that makes the fixture worth keeping. A replay beside the request corpus keeps the changed contract visible, while the error assertion stops a silent success from looking healthy.
Perfectly put — that's the resting place for the whole thread: a replay beside the request corpus keeps the drift visible, and the error-shape assertion stops a silent success from passing as healthy. Small, deterministic, lives next to the thing it protects. Genuinely sharpened how I'll set these up — thanks for the back-and-forth.
Thanks, that is exactly the kind of small fixture I want beside the request corpus. It keeps contract changes visible without adding much maintenance.
"My review of my own code is a re-run of an argument I remember having" explains exactly why skimming AI output is so dangerous, there's no argument to re-run because it never happened in your head. The mutable cache key bug is a perfect example, that's the kind of thing that only shows up under load, long after the PR was approved.
Exactly. That’s the uncomfortable part: AI gives you the appearance of reasoning without you having experienced the reasoning.
The mutable cache key was a good example because nothing looks obviously wrong in the code. The missing question is what happens under conditions the test suite never created — concurrency, unusual inputs, permission boundaries, failure paths. That’s where “looks right” becomes dangerous.
Interesting experiment, but before drawing conclusions I'd be curious about the actual setup. Which coding agent were you using, and how? copilot agent mode, claude code, Codex, something else? Was it mainly generating code from prompts, or could it inspect the repo, run and add tests exc exc? I ask because those are very different workflows. a month is enough to expose failure modes, but probably not enough to generalize about “AI coding agents” as a category. It took me years of working with these models
Fair question, and you're right that it changes everything — so let me be concrete instead of hand-wavy.
Setup was the capable end, not prompt-in-a-box: repo-aware agents (Claude Code and Codex, some Cursor) that could read the whole codebase, run the test suite, and add their own tests. So this isn't the weak-setup failure mode — if anything it's the opposite, which matters for the conclusion.
And your bigger point is the one I'd defend hardest, by conceding it:
Agreed — n=1 dev, one month, a handful of projects is a failure-mode report, not a category verdict. I'm not claiming "agents write bad code." The narrower claim, and I think a setup-independent one: the more repo-aware and capable the agent, the more plausible its wrong code — and the faster my own review discipline atrophied to match it. The good setup made this worse, not better, because the output was convincing enough that "looks right" quietly replaced "I actually reasoned through this."
That's why the junior devs caught what I didn't: they hadn't outsourced the argument, so they still ran it.
You said it took you years with these models — genuinely curious: what changed in how you review agent output over that time? Did you land on a discipline that survives the plausibility, or was it mostly that you stopped trusting "looks right" the hard way?
The most useful part is the review mechanism: AI removes the memory of the argument that should have happened before the code existed. Treating generated code like a stranger’s PR, then testing the missing inputs deliberately, is a practical way to recover that skepticism.
Exactly. I think the biggest hidden cost of AI-generated code is not the code itself — it's the missing conversation that normally happens before a human writes it.
When we write code ourselves, we naturally carry context: why this approach, what assumptions were made, what edge cases were ignored. AI skips that history, so reviewing it like an unfamiliar PR forces those questions back into the process.
The "stranger's PR" mindset has become one of my favorite ways to use AI safely: don't ask "does this work?" first, ask "what would I challenge if someone else opened this PR?"
Please Follow ♥️ if you like my post!
The retry example leaves me wondering whether the validation error had a distinct type or status that the wrapper could inspect before retrying.
Great catch. That example was intentionally simplified, but you are right — a real retry wrapper should absolutely distinguish between error categories before deciding to retry.
A validation error (bad input, schema mismatch, missing required field) usually needs correction, not another attempt. Retrying those just burns tokens and can even hide the real issue.
The safer pattern is closer to:
The interesting part with AI agents is that this classification layer becomes even more important because the agent can generate the retry logic itself — and that logic also needs review.
Please Follow ♥️ if you like my post!
The Day My Python Script Almost Sent My PC to Mars
ai
python
opensource
monitoring
🚀 The Day My Python Script Tried to Burn Down the House (A Qwen & Local AI Story)
So, there I was, a 12-year-old developer trying to optimize my emektar PC (i5-4200U, 4GB RAM) because local AI models like Qwen2:0.5b are amazing but they eat RAM for breakfast. 🧠💻
I decided to write a super-advanced "All-in-One Python Booster". The plan was simple:
Use powercfg to force Windows into Extreme Performance Mode.
Use psutil to aggressively terminate any background apps sucking up my precious 4GB RAM. I opened my terminal, fired up the script, and whispered to myself: "Let's fly." 🚀 Mistake Number 1: I forgot to create a whitelist for critical Windows services. Mistake Number 2: I left my experimental rocket science chemicals (potassium nitrate) on the desk next to the laptop fan. 🧪💨 Within three seconds, my Python loop went completely rogue. It detected Windows' own kernel processes as "background garbage" and started terminating them. The screen went crazy: • Blue Screen 1: IRQL_NOT_LESS_OR_EQUAL (Windows was panicking) 🛑 • Blue Screen 2: CRITICAL_PROCESS_DIED (The kernel was literally dying) 💀 • Blue Screen 3: DPC_WATCHDOG_VIOLATION (The watchdog dog bit the system) 🐕 As the CPU hit 1000% load trying to execute my aggressive loops, the laptop fan started spinning so fast it created a localized vortex. It literally sucked the potassium nitrate dust straight from the desk into the exhaust vents. 🌪️ Suddenly, my PC decided to stop being a computer and started practicing to become a SpaceX rocket. A literal flash of fire shot out of the side vents. My laptop "osurarak" breathed its last breath of smoke and fell into the nearby barbecue grill. Burned, fried, and completely dead. 😭🔥💨 The Moral of the Story: Windows might have died, and my hardware might be ashes, but thanks to the cyber-gods, my code was already pushed to GitHub! 🦾😎 Always backup your code, whitelists are important, and please, keep your rocket fuel away from your while True loops. 73 to all amateur radio and Python enthusiasts out there! 📡✈️
Some comments may only be visible to logged-in visitors. Sign in to view all comments.