DEV Community

Cover image for I let AI write my code for a month. The junior devs caught what it broke — I didn't.

I let AI write my code for a month. The junior devs caught what it broke — I didn't.

Rudratosh Shastri on September 29, 2026

I'll say the unpopular part first: for one month I leaned on an AI coding agent for almost everything, and the bugs it introduced were caught by th...
Collapse
 
vladzoff profile image
Vlad Zoff •

The point about tests only covering the questions you already wrote down has another version at the architecture level. I've seen agent changes keep every test green while quietly bypassing an existing module boundary or adding a dependency in the wrong direction. That's harder to notice because the behavior is still correct.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

agent changes keep every test green while quietly bypassing an existing module boundary or adding a dependency in the wrong direction

This is the scarier sibling of the point, and you've named it precisely: the behavior is correct, so no test fires, but the architecture quietly rots. Tests assert behavior; they almost never assert structure.

The only thing that's caught this for me is making structure itself testable — dependency-direction lint / import boundaries (dependency-cruiser, an ArchUnit-style check) — so "UI imported from infra" fails CI even when the feature works. An agent can't argue with a boundary encoded as a rule it can see.

Have you found a lint/arch check that catches the "right direction, wrong dependency" case, or is it still human review doing the catching?

Collapse
 
vladzoff profile image
Vlad Zoff •

Yes, dependency-direction checks are a good first layer. The harder case is exactly the one you mentioned: the dependency follows the allowed direction, but the specific dependency is still wrong for that boundary.
For example, route -> service may be valid, while one particular route should not depend on a specific service, or a service may be allowed to use a repository but not bypass another application boundary. That's difficult to express as a generic import-direction rule because the answer depends on the project's actual architecture.
That's the part I'm working on (with Guard): discover the existing structure first, then let the project define those more specific boundaries as rules/contracts. Human review still makes sense for cases where the question is really about intent rather than something the repository can express deterministically.

Thread Thread
 
rudratosh profile image
Rudratosh Shastri •

the dependency follows the allowed direction, but the specific dependency is still wrong for that boundary

That's the gap generic lint can't close. Direction rules encode the layering, not the architecture. "Routes may call services" says nothing about this route calling that service.

Discovering the existing structure first and then letting the project declare contracts sounds like the right order, because:

  • Hand-written rules from scratch never get written
  • Inferred rules with no human sign-off will happily bless the existing mistakes
  • Inferred, then confirmed gives you a baseline someone actually agreed to

That second point is my real question. When Guard discovers the current structure, how does it tell an intentional boundary from an accident that just happens to be consistent so far?

Collapse
 
elijahbrown profile image
Elijah Brown •

The plausible-bug point lands. One case I'd add to 'what input is not in these tests' for AI-written forms: a well-formed email on a domain with no MX, and a phone typed in local format for the wrong country. Format checks pass both; a server-side parse catches them.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

Exactly — those are great examples of the “looks valid, behaves wrong” class of bugs I was getting at.

A syntactically valid email isn't necessarily deliverable, and a correctly formatted phone number isn't necessarily valid for the user's country. Those are precisely the cases that a happy-path test suite can miss because the input satisfies the validator while violating the real-world constraint.

And the server-side validation point is important too: the client can check shape, but the backend needs to validate the actual semantics before trusting the value. That's another good example of why “the tests pass” isn't the same as “the input space is covered.”

Collapse
 
elijahbrown profile image
Elijah Brown •

Agreed, and both make cheap test fixtures: a domain that publishes a Null MX, and a national-format number sent with no country, since the same digits can be valid in one country and invalid in another.

Thread Thread
 
rudratosh profile image
Rudratosh Shastri •

Exactly. Those are great examples because neither requires an exotic edge case — they’re cheap, deterministic fixtures that expose whether the implementation actually understands the semantics.

That’s the bigger lesson with AI-generated code: a handful of deliberately adversarial test cases can reveal gaps that a large pile of happy-path tests completely misses.

Thread Thread
 
elijahbrown profile image
Elijah Brown •

Exactly. Small deterministic cases like those are a good way to expose semantic gaps that happy-path tests miss.

Thread Thread
 
rudratosh profile image
Rudratosh Shastri •

Small deterministic cases … expose semantic gaps that happy-path tests miss

Exactly — and the deterministic part is what makes them cheap to keep: a Null-MX domain and a national-format number with no country code never go stale, never flake, and each encodes a real-world constraint the validator doesn't. I've started keeping a little "adversarial fixtures" file that's just these — the cases I'd never have written if I hadn't been burned. What's the one you reach for most?

Collapse
 
respect17 profile image
Kudzai Murimi •

"My review of my own code is a re-run of an argument I remember having" explains exactly why skimming AI output is so dangerous, there's no argument to re-run because it never happened in your head. The mutable cache key bug is a perfect example, that's the kind of thing that only shows up under load, long after the PR was approved.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

Exactly. That’s the uncomfortable part: AI gives you the appearance of reasoning without you having experienced the reasoning.

The mutable cache key was a good example because nothing looks obviously wrong in the code. The missing question is what happens under conditions the test suite never created — concurrency, unusual inputs, permission boundaries, failure paths. That’s where “looks right” becomes dangerous.

Collapse
 
danielecangi profile image
DaC •

Interesting experiment, but before drawing conclusions I'd be curious about the actual setup. Which coding agent were you using, and how? copilot agent mode, claude code, Codex, something else? Was it mainly generating code from prompts, or could it inspect the repo, run and add tests exc exc? I ask because those are very different workflows. a month is enough to expose failure modes, but probably not enough to generalize about “AI coding agents” as a category. It took me years of working with these models

Collapse
 
rudratosh profile image
Rudratosh Shastri •

before drawing conclusions I'd be curious about the actual setup. Which coding agent were you using, and how?

Fair question, and you're right that it changes everything — so let me be concrete instead of hand-wavy.

Setup was the capable end, not prompt-in-a-box: repo-aware agents (Claude Code and Codex, some Cursor) that could read the whole codebase, run the test suite, and add their own tests. So this isn't the weak-setup failure mode — if anything it's the opposite, which matters for the conclusion.

And your bigger point is the one I'd defend hardest, by conceding it:

a month is enough to expose failure modes, but probably not enough to generalize about "AI coding agents" as a category

Agreed — n=1 dev, one month, a handful of projects is a failure-mode report, not a category verdict. I'm not claiming "agents write bad code." The narrower claim, and I think a setup-independent one: the more repo-aware and capable the agent, the more plausible its wrong code — and the faster my own review discipline atrophied to match it. The good setup made this worse, not better, because the output was convincing enough that "looks right" quietly replaced "I actually reasoned through this."

That's why the junior devs caught what I didn't: they hadn't outsourced the argument, so they still ran it.

You said it took you years with these models — genuinely curious: what changed in how you review agent output over that time? Did you land on a discipline that survives the plausibility, or was it mostly that you stopped trusting "looks right" the hard way?

Collapse
 
maximin_mxn_6ce1a3054be6d profile image
Maximin •

The most useful part is the review mechanism: AI removes the memory of the argument that should have happened before the code existed. Treating generated code like a stranger’s PR, then testing the missing inputs deliberately, is a practical way to recover that skepticism.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

Exactly. I think the biggest hidden cost of AI-generated code is not the code itself — it's the missing conversation that normally happens before a human writes it.

When we write code ourselves, we naturally carry context: why this approach, what assumptions were made, what edge cases were ignored. AI skips that history, so reviewing it like an unfamiliar PR forces those questions back into the process.

The "stranger's PR" mindset has become one of my favorite ways to use AI safely: don't ask "does this work?" first, ask "what would I challenge if someone else opened this PR?"

Please Follow ♥️ if you like my post!

Collapse
 
kyisaiah47 profile image
kyisaiah47 •

The retry example leaves me wondering whether the validation error had a distinct type or status that the wrapper could inspect before retrying.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

Great catch. That example was intentionally simplified, but you are right — a real retry wrapper should absolutely distinguish between error categories before deciding to retry.

A validation error (bad input, schema mismatch, missing required field) usually needs correction, not another attempt. Retrying those just burns tokens and can even hide the real issue.

The safer pattern is closer to:

  • retry transient failures (timeouts, rate limits, temporary upstream errors)
  • stop and surface deterministic failures (validation, permission, business rules)
  • log the context that led to the failure

The interesting part with AI agents is that this classification layer becomes even more important because the agent can generate the retry logic itself — and that logic also needs review.

Please Follow ♥️ if you like my post!

Collapse
 
kozmonot20 profile image
kozmonot20 •

The Day My Python Script Almost Sent My PC to Mars

ai

python

opensource

monitoring
🚀 The Day My Python Script Tried to Burn Down the House (A Qwen & Local AI Story)
So, there I was, a 12-year-old developer trying to optimize my emektar PC (i5-4200U, 4GB RAM) because local AI models like Qwen2:0.5b are amazing but they eat RAM for breakfast. 🧠💻
I decided to write a super-advanced "All-in-One Python Booster". The plan was simple:

Use powercfg to force Windows into Extreme Performance Mode.
Use psutil to aggressively terminate any background apps sucking up my precious 4GB RAM. I opened my terminal, fired up the script, and whispered to myself: "Let's fly." 🚀 Mistake Number 1: I forgot to create a whitelist for critical Windows services. Mistake Number 2: I left my experimental rocket science chemicals (potassium nitrate) on the desk next to the laptop fan. 🧪💨 Within three seconds, my Python loop went completely rogue. It detected Windows' own kernel processes as "background garbage" and started terminating them. The screen went crazy: • Blue Screen 1: IRQL_NOT_LESS_OR_EQUAL (Windows was panicking) 🛑 • Blue Screen 2: CRITICAL_PROCESS_DIED (The kernel was literally dying) 💀 • Blue Screen 3: DPC_WATCHDOG_VIOLATION (The watchdog dog bit the system) 🐕 As the CPU hit 1000% load trying to execute my aggressive loops, the laptop fan started spinning so fast it created a localized vortex. It literally sucked the potassium nitrate dust straight from the desk into the exhaust vents. 🌪️ Suddenly, my PC decided to stop being a computer and started practicing to become a SpaceX rocket. A literal flash of fire shot out of the side vents. My laptop "osurarak" breathed its last breath of smoke and fell into the nearby barbecue grill. Burned, fried, and completely dead. 😭🔥💨 The Moral of the Story: Windows might have died, and my hardware might be ashes, but thanks to the cyber-gods, my code was already pushed to GitHub! 🦾😎 Always backup your code, whitelists are important, and please, keep your rocket fuel away from your while True loops. 73 to all amateur radio and Python enthusiasts out there! 📡✈️