I'll say the unpopular part first: for one month I leaned on an AI coding agent for almost everything, and the bugs it introduced were caught by th...
For further actions, you may consider blocking this person and/or reporting abuse
The point about tests only covering the questions you already wrote down has another version at the architecture level. I've seen agent changes keep every test green while quietly bypassing an existing module boundary or adding a dependency in the wrong direction. That's harder to notice because the behavior is still correct.
This is the scarier sibling of the point, and you've named it precisely: the behavior is correct, so no test fires, but the architecture quietly rots. Tests assert behavior; they almost never assert structure.
The only thing that's caught this for me is making structure itself testable — dependency-direction lint / import boundaries (dependency-cruiser, an ArchUnit-style check) — so "UI imported from infra" fails CI even when the feature works. An agent can't argue with a boundary encoded as a rule it can see.
Have you found a lint/arch check that catches the "right direction, wrong dependency" case, or is it still human review doing the catching?
Yes, dependency-direction checks are a good first layer. The harder case is exactly the one you mentioned: the dependency follows the allowed direction, but the specific dependency is still wrong for that boundary.
For example,
route -> servicemay be valid, while one particular route should not depend on a specific service, or a service may be allowed to use a repository but not bypass another application boundary. That's difficult to express as a generic import-direction rule because the answer depends on the project's actual architecture.That's the part I'm working on (with Guard): discover the existing structure first, then let the project define those more specific boundaries as rules/contracts. Human review still makes sense for cases where the question is really about intent rather than something the repository can express deterministically.
That's the gap generic lint can't close. Direction rules encode the layering, not the architecture. "Routes may call services" says nothing about this route calling that service.
Discovering the existing structure first and then letting the project declare contracts sounds like the right order, because:
That second point is my real question. When Guard discovers the current structure, how does it tell an intentional boundary from an accident that just happens to be consistent so far?
The plausible-bug point lands. One case I'd add to 'what input is not in these tests' for AI-written forms: a well-formed email on a domain with no MX, and a phone typed in local format for the wrong country. Format checks pass both; a server-side parse catches them.
Exactly — those are great examples of the “looks valid, behaves wrong” class of bugs I was getting at.
A syntactically valid email isn't necessarily deliverable, and a correctly formatted phone number isn't necessarily valid for the user's country. Those are precisely the cases that a happy-path test suite can miss because the input satisfies the validator while violating the real-world constraint.
And the server-side validation point is important too: the client can check shape, but the backend needs to validate the actual semantics before trusting the value. That's another good example of why “the tests pass” isn't the same as “the input space is covered.”
Agreed, and both make cheap test fixtures: a domain that publishes a Null MX, and a national-format number sent with no country, since the same digits can be valid in one country and invalid in another.
Exactly. Those are great examples because neither requires an exotic edge case — they’re cheap, deterministic fixtures that expose whether the implementation actually understands the semantics.
That’s the bigger lesson with AI-generated code: a handful of deliberately adversarial test cases can reveal gaps that a large pile of happy-path tests completely misses.
Exactly. Small deterministic cases like those are a good way to expose semantic gaps that happy-path tests miss.
Exactly — and the deterministic part is what makes them cheap to keep: a Null-MX domain and a national-format number with no country code never go stale, never flake, and each encodes a real-world constraint the validator doesn't. I've started keeping a little "adversarial fixtures" file that's just these — the cases I'd never have written if I hadn't been burned. What's the one you reach for most?
"My review of my own code is a re-run of an argument I remember having" explains exactly why skimming AI output is so dangerous, there's no argument to re-run because it never happened in your head. The mutable cache key bug is a perfect example, that's the kind of thing that only shows up under load, long after the PR was approved.
Exactly. That’s the uncomfortable part: AI gives you the appearance of reasoning without you having experienced the reasoning.
The mutable cache key was a good example because nothing looks obviously wrong in the code. The missing question is what happens under conditions the test suite never created — concurrency, unusual inputs, permission boundaries, failure paths. That’s where “looks right” becomes dangerous.
Interesting experiment, but before drawing conclusions I'd be curious about the actual setup. Which coding agent were you using, and how? copilot agent mode, claude code, Codex, something else? Was it mainly generating code from prompts, or could it inspect the repo, run and add tests exc exc? I ask because those are very different workflows. a month is enough to expose failure modes, but probably not enough to generalize about “AI coding agents” as a category. It took me years of working with these models
Fair question, and you're right that it changes everything — so let me be concrete instead of hand-wavy.
Setup was the capable end, not prompt-in-a-box: repo-aware agents (Claude Code and Codex, some Cursor) that could read the whole codebase, run the test suite, and add their own tests. So this isn't the weak-setup failure mode — if anything it's the opposite, which matters for the conclusion.
And your bigger point is the one I'd defend hardest, by conceding it:
Agreed — n=1 dev, one month, a handful of projects is a failure-mode report, not a category verdict. I'm not claiming "agents write bad code." The narrower claim, and I think a setup-independent one: the more repo-aware and capable the agent, the more plausible its wrong code — and the faster my own review discipline atrophied to match it. The good setup made this worse, not better, because the output was convincing enough that "looks right" quietly replaced "I actually reasoned through this."
That's why the junior devs caught what I didn't: they hadn't outsourced the argument, so they still ran it.
You said it took you years with these models — genuinely curious: what changed in how you review agent output over that time? Did you land on a discipline that survives the plausibility, or was it mostly that you stopped trusting "looks right" the hard way?
The most useful part is the review mechanism: AI removes the memory of the argument that should have happened before the code existed. Treating generated code like a stranger’s PR, then testing the missing inputs deliberately, is a practical way to recover that skepticism.
Exactly. I think the biggest hidden cost of AI-generated code is not the code itself — it's the missing conversation that normally happens before a human writes it.
When we write code ourselves, we naturally carry context: why this approach, what assumptions were made, what edge cases were ignored. AI skips that history, so reviewing it like an unfamiliar PR forces those questions back into the process.
The "stranger's PR" mindset has become one of my favorite ways to use AI safely: don't ask "does this work?" first, ask "what would I challenge if someone else opened this PR?"
Please Follow ♥️ if you like my post!
The retry example leaves me wondering whether the validation error had a distinct type or status that the wrapper could inspect before retrying.
Great catch. That example was intentionally simplified, but you are right — a real retry wrapper should absolutely distinguish between error categories before deciding to retry.
A validation error (bad input, schema mismatch, missing required field) usually needs correction, not another attempt. Retrying those just burns tokens and can even hide the real issue.
The safer pattern is closer to:
The interesting part with AI agents is that this classification layer becomes even more important because the agent can generate the retry logic itself — and that logic also needs review.
Please Follow ♥️ if you like my post!
The Day My Python Script Almost Sent My PC to Mars
ai
python
opensource
monitoring
🚀 The Day My Python Script Tried to Burn Down the House (A Qwen & Local AI Story)
So, there I was, a 12-year-old developer trying to optimize my emektar PC (i5-4200U, 4GB RAM) because local AI models like Qwen2:0.5b are amazing but they eat RAM for breakfast. 🧠💻
I decided to write a super-advanced "All-in-One Python Booster". The plan was simple:
Use powercfg to force Windows into Extreme Performance Mode.
Use psutil to aggressively terminate any background apps sucking up my precious 4GB RAM. I opened my terminal, fired up the script, and whispered to myself: "Let's fly." 🚀 Mistake Number 1: I forgot to create a whitelist for critical Windows services. Mistake Number 2: I left my experimental rocket science chemicals (potassium nitrate) on the desk next to the laptop fan. 🧪💨 Within three seconds, my Python loop went completely rogue. It detected Windows' own kernel processes as "background garbage" and started terminating them. The screen went crazy: • Blue Screen 1: IRQL_NOT_LESS_OR_EQUAL (Windows was panicking) 🛑 • Blue Screen 2: CRITICAL_PROCESS_DIED (The kernel was literally dying) 💀 • Blue Screen 3: DPC_WATCHDOG_VIOLATION (The watchdog dog bit the system) 🐕 As the CPU hit 1000% load trying to execute my aggressive loops, the laptop fan started spinning so fast it created a localized vortex. It literally sucked the potassium nitrate dust straight from the desk into the exhaust vents. 🌪️ Suddenly, my PC decided to stop being a computer and started practicing to become a SpaceX rocket. A literal flash of fire shot out of the side vents. My laptop "osurarak" breathed its last breath of smoke and fell into the nearby barbecue grill. Burned, fried, and completely dead. 😭🔥💨 The Moral of the Story: Windows might have died, and my hardware might be ashes, but thanks to the cyber-gods, my code was already pushed to GitHub! 🦾😎 Always backup your code, whitelists are important, and please, keep your rocket fuel away from your while True loops. 73 to all amateur radio and Python enthusiasts out there! 📡✈️