So I did the thing everyone tells you not to do.
I took the whole codebase. Every route, the migrations, the config, that utils.ts nobody's touched since 2023. Dumped all of it into a model with a context window big enough to swallow the entire repo in one go, and asked it one question:
"What's wrong with this?"
Now, everyone's first reaction to that is the security angle. "You pasted your source into a chatbot?" Fair. I'll get to it, and yeah, it matters. But that's not the part that got to me.
What got to me was the answer. It came back as a clean numbered list, and I sat there realizing I couldn't actually tell which items were true and which weren't. They all read the same.
Let me walk through it the way it happened.
First it impressed the hell out of me
I expected file summaries. What I got was my architecture, described back to me. It knew auth ran through one middleware. It knew two of my services were quietly sharing a table they had no business sharing. It knew the payment webhook and the signup path both wrote to users from totally different places in the code.
Honestly? It mapped the system better than the onboarding doc I'd written for new hires. Faster, too.
If you've never tried this, do it once just for this. A model sitting on your entire repo can answer "where does X happen, and what breaks if I change it" better than half the people who actually work in the thing. That part is real, and it's worth your time.
Keep that in mind though, because it's exactly what set me up.
Then it found bugs. Real ones.
This is where I got excited. It started surfacing actual problems:
- An old
.envI'd committed back in 2022 and deleted the same week. Still sitting in git history, of course. Forever. I knew that in the abstract. I had completely forgotten it in practice. - A race between the webhook handler and the signup path. Both upsert the same row, no lock. Under a retry, one clobbers the other. Nobody had hit it yet, but it was absolutely there.
- A dead admin endpoint. Unhooked from the UI two years ago, still mounted, still skipping the newer permission check. Reachable. Forgotten.
These weren't lint warnings. These were "how did I ship that" bugs. For a solid ten minutes I thought I was writing a very different post, the one titled "just give ChatGPT your codebase already, it's incredible."
And look, as a finder, it kind of is incredible. It's read more code than I ever will and it doesn't get tired on file number 340. I'm not going to pretend otherwise.
Then it started making things up. In the exact same voice.
Item 7 was a SQL injection in one of my functions.
That function doesn't exist.
It had everything going for it, though. Plausible name, plausible file, a nice confident description of how an attacker would exploit it, the whole "this could lead to data exfiltration" bit. Every signal your brain uses to go "ok this person knows what they're talking about" was firing. The only thing missing was the code it was describing.
Item 9 flagged a missing auth check on an endpoint that had the check. It was one file over from where it was looking.
Here's what actually rattled me. Items 7 and 9 looked identical to items 1 through 3. Same confidence. Same layout. Same tidy little "here's how to fix it" block. Nothing in the way it wrote the fake ones told me they were fake.
More context hadn't made it more honest. It just gave it more of my real function names to wrap a made-up story around. And a hallucination wearing your own variable names is a lot more convincing than a generic one, trust me.
And then the one that actually scared me
Buried in the webhook handler was the real landmine. Not the race this time. Something quieter:
// payment webhook
async function handle(event) {
ack(event); // tell the provider "got it" → 200
await saveToDb(event); // ...then try to persist
}
The ack goes out before the write is safe. If the process dies, or the DB throws, in that tiny window after the acknowledgement, the provider thinks it's done and never retries. The payment is just gone. Silently. And every log line is green. (If you've read the 2am outage post, you already know how this movie ends.)
This is the single most dangerous thing in the repo. Data loss with no error to point at.
So I asked it straight up: "anything wrong with the webhook handler?"
It told me the handler looked correct. Said the try/catch was good practice. Suggested I add a comment.
Same tone it had just used to correctly nuke my dead admin endpoint. Same tone it used to invent a SQL injection out of thin air. On the one bug that can actually lose a customer's money, it gave me a thumbs up and a note about code style.
And I get why. The bug isn't really in the handler. It's in the relationship between two lines, plus a fact that lives outside my codebase entirely: how the payment provider handles retries. That's a seam bug. The model reads inside the frame you hand it, and the danger was in the frame, not the picture.
The real problem: I couldn't trust my own read anymore
Step back and look at what I was actually holding:
- 3 findings that were true and genuinely useful
- 2 that were confidently wrong
- 1 catastrophic bug it signed off on
And nothing in the output told me which was which. The fake vulnerability read exactly as credible as the real one. The "this is fine" on the webhook read exactly as credible as a "this is fine" on genuinely clean code would have.
That's the scary part. Not "AI writes bugs." Not even "AI hallucinates," we all know that by now. It's this:
A model sitting on your whole codebase gives you output where the confidence has basically nothing to do with whether it's right. And more context cranks up the confidence without doing anything for the bugs that live between your files.
I wanted a reviewer. What I got was the most convincing narrator of my own code I've ever seen. Equally convincing when it was right, when it was wrong, and when it was about to cost me money.
Why more context makes this worse, not better
You'd assume more context = better review. For finding stuff, sure. For judgement, it actually cuts the wrong way:
More surface area to sound expert about. It can now drop your real module names into a completely fabricated claim, and specificity reads as truth. The seam bugs are still invisible, because whole-repo context doesn't include the world your repo runs in, and that's where the expensive bugs hide. And it's reviewing with the same instincts it would've used to write the thing, so right where it would've made a mistake, it can't see the mistake, because catching it would mean disagreeing with itself.
None of this is "don't use it." It's "know which job you're actually asking it to do."
Finder vs. witness
There are two different jobs hiding inside the word "review," and I'd been treating them as one.
A finder throws candidate problems at you. Recall is what matters. Being wrong sometimes is completely fine, because a human triages the pile. A whole-repo model is a great finder. Use it as one, no notes.
A witness is the thing that says "yep, this is correct, ship it." And here, being confidently wrong is a disaster, because the entire point of a witness is that nobody checks its work again.
The model is an A+ finder and an F witness, and it hands you both in the same paragraph without flagging the difference. Every mistake I made that afternoon came from reading its witness statements like they carried a finder's stakes.
So here's what I do now:
- Treat it as a finder, never a witness. Every "this looks fine" is worth zero to me. Only the "here's a problem" items get my attention, and each one is a lead, not a verdict.
- Check every finding against the actual code. The two hallucinations died in about thirty seconds once I opened the file. That check costs nothing. Skipping it costs you a day "fixing" a bug that was never there, feeling productive the whole time.
- Whatever certifies the change has to be independent. Different prior than the thing that wrote it. A different model family, an adversarial prompt ("give me the input that loses money" beats "is this correct?"), a test that actually runs, and a human on the merge button. A second pass from the same mind just re-derives the same blind spot.
- Seam bugs need seam tests. No amount of reading catches ack-before-persist. Kill the process between those two lines and watch what the provider does. Behaviour, not opinion.
Right, the security part
Pasting a proprietary codebase into a consumer chat product is a decision, not a reflex. Before you do what I did:
- Assume anything you put in a consumer tier might be retained or trained on unless a contract says otherwise. Read the terms for the exact tier you're on, not the marketing page.
- Secrets are the immediate risk. The model found a secret in my git history, which means the secret was in what I pasted. Scrub credentials, tokens, customer data, internal hostnames before anything leaves your machine.
- If this is real work, use an enterprise tier with a no-training guarantee, or run a model where your code already lives. The convenience of the chat box is not worth your source tree showing up in a training set later.
I did the whole thing on a throwaway clone with the secrets already rotated. Do that.
So, the actual takeaway
Give a model your whole codebase. Seriously, do it. As a finder it'll show you stuff you shipped and forgot, and it'll map your system faster than your own docs.
Just don't call that a review. The thing I really walked away with is that its "this is correct" isn't evidence of anything, and it shows up in the exact same voice as the findings that are pure gold. The second you let the thing that reads the code also be the thing that certifies the code, you've built yourself a witness that agrees with itself every single time.
A model saying "looks correct" was never proof it's correct. You still need a second seat whose only job is to not believe the first one. And a human on the merge.
I build xenition on exactly that split: one model writes, a separate skeptic with a different prior tries to tear it apart and isn't allowed to say "looks fine," and a human owns the merge. This whole mess is why.
Top comments (13)
The interesting part isn’t that AI can understand a large codebase — it’s how quickly it can expose assumptions and technical debt that humans have gotten used to.
The scary part is realizing how much “context” lives in our heads instead of the code itself. 😅
Great read.
That line about context living in our heads is the one, yeah. The AI mapped my architecture better than my own onboarding doc — but only the part that was written down. Everything that made the system actually safe to change lived nowhere: "don't touch that table directly," "the webhook can fire twice," "this endpoint looks dead but a cron still hits it." None of that is in the code. It's in whoever's been here longest.
And here's the twist that got me — the model doesn't know that context is missing. It reads the repo, sees no note saying "careful here," and concludes there's nothing to be careful about. So the exact spots where our head-knowledge was load-bearing are the spots it signs off on most confidently. The silence in the code reads to it as "all clear."
Which is kind of a brutal mirror, honestly. Every "looks fine" it hands you is really measuring how much of your own understanding you forgot to write down. The debt was never just in the code — it's in the gap between the code and the stuff we all just know.
Appreciate you reading it.
The runtime half of that gap makes me think of my opposite case. My tools are all static sites with no backend — no async chains, no distributed system — so there's nothing here to crash-simulate. But I do use runtime probes every day. I just didn't write them.
I hand my production sites to two outside readers as free probes. One is search crawlers: after each page ships, I check Search Console's URL Inspection for how it actually reads the page and whether it got indexed. My newest site's first post was indexed in under 24 hours — faster feedback than any local check. The other is PageSpeed: it loads your page from real edge nodes. It once caught a 177 KB third-party script choking my homepage, and five mobile scores moved from the 80s to near 100. Static reading barely sees that problem, because it only exists under real loading.
What both probes share: they look at the real-world you, not a mirror of your code. AI reading code is still reading what you wrote. Crawlers and speed tools report what you shipped.
So for a small site with no backend, the runtime half doesn't need a homemade simulation. Hand it to the outside readers who are already reading you.
This is the best extension of the argument anyone's left, honestly. You found the witness I kept insisting had to be a human or a separate model, and it turns out for a static site it already exists and you don't even have to build it.
The thing I love about your two probes is that they both satisfy the one rule that actually matters: independent prior. Search Console and PageSpeed aren't reading your intent, they're reporting your consequences. The crawler doesn't care what your HTML meant to do, it tells you what it actually did when a real indexer hit it. That's exactly the "different seat" I was groping toward — you just realized the seat was already occupied by Google and you were ignoring the guy sitting in it.
And the 177 KB script is the perfect example of a seam bug in your world. Statically, that script is fine — correct tag, loads, no error. The bug only exists in the relationship between your page and a real edge node on a real phone on a real network. No amount of reading the source surfaces it, same way no amount of reading my webhook surfaced the ack-before-persist. The danger was in the frame, not the picture — your frame's just the network instead of a payment provider's retry logic.
The one place I'd gently push: those probes are fantastic witnesses for the dimensions they measure — indexability, load behavior. They're silent on correctness of the thing itself. A page can be indexed in 24 hours, score 100 on mobile, and still show the wrong price. So I'd say you've fully solved the runtime half for performance and reachability, and the "did the content actually come out right" half still wants a human eye or a test. But for the half you're talking about? Yeah — don't simulate it, just go read what your outside readers already wrote about you. That's genuinely sharper than what I said in the post.
That distinction between a finder and a witness is easily the sharpest takeaway here.
The scariest hallucinations are never the obvious syntax crashes. They are the ones dressed in your own real architecture, using your actual function names, explaining a phantom vulnerability with total executive poise. Worse still is the reverse: an AI giving a green checkmark to an ack-before-persist pattern just because the try/catch block looks polite. It understands your syntax, but it has zero clue how the external world behaves when a socket drops mid-flight.
Treating every positive output as a lead instead of a verdict is the only sane approach.
What is the most convincing, beautiful-looking lie an AI has told you about your own codebase?
You nailed the reverse case — "the try/catch looks polite" is exactly the tone. It grades the manners of the code and calls it a security review.
The most convincing lie I ever got: it told me I had a timing-attack vulnerability in my password comparison. Named the function. Quoted the line where I was supposedly using === on the hash instead of a constant-time compare. Walked me through how an attacker measures response times to recover the hash byte by byte. Textbook, genuinely well-written, the kind of finding that makes you feel like you dodged a real bullet.
I was constant-time-comparing. Had been the whole time. It quoted a version of my function that didn't exist — right file, right function name, wrong body. It had basically written the bug it expected to find there and handed it back to me as something it found.
And that's the tell, I think, looking back: it didn't read my code and report a problem. It pattern-matched "password comparison" to "probably a timing attack" and then dressed the prediction up in my real function name. The specificity I trusted — my name, my file — was the cheapest part for it to generate. The one part that would've taken actually reading the code, the function body, was the part it made up.
Thirty seconds in the file killed it. But I'll be honest, for those thirty seconds I believed it completely. That's the gap that scares me — not that it lied, but that nothing in the lie felt different from the three findings right above it that were real.
What you said is a real story. Once AI finishes reading a codebase, it can genuinely flag issues you never noticed — some pretty serious ones, too. But that doesn’t make it rigorous or trustworthy itself; it’ll happily write serious bugs of its own. Tests have to be even stricter than before.
Yeah, exactly. And honestly that's the part that still gets me: the same tool that catches a bug you'd never have spotted will hand you a confident, perfectly-formatted bug two lines later, and on the page they look identical.
Fully agree on stricter tests. One thing I'd add though: if the AI writes the tests too, they inherit its blind spot. It'll happily test the behaviour it already believes is correct. So "stricter" has to also mean independent — tests from a different prior, plus the behavioural/seam ones for bugs that don't live inside a single file (the webhook ack-before-save never shows up in a normal unit test). Otherwise you just end up with greener tests that all quietly agree with each other.
"The model reads inside the frame you hand it, and the danger was in the frame, not the picture."
That line completely nails it.
LLMs analyze static text, but production outages happen in distributed state transitions. You can give a model your entire codebase, but you can't give it the network partitions, third-party webhook retry backoffs, or the database pool running out of connections at 3 AM.
Static code review (human or AI) verifies syntax and obvious logic. Fault injection and integration testing verify reality. If you want to catch seam bugs, stop asking models to read the code—start simulating process crashes between async lines.
This is the sharper version of what I was clumsily getting at, so thank you.
"Stop asking models to read the code, start simulating crashes between async lines" is exactly it. The one place I'd still keep the model in the loop is finding where to inject. Point it at the repo and ask "list every spot where we ack, respond, or commit before the durable write, or hold a DB connection across an await," and it's genuinely good at enumerating the candidate seams. Then you fault-inject each one and let reality vote.
Model finds the suspects, chaos testing convicts them. It just never gets to be the judge.
I especially like the point that giving AI a full codebase can reveal problems you’ve simply stopped noticing after working with the same project for a long time.
Yeah exactly. after a while you just stop seeing certain things. that dead admin endpoint and the old .env in git history were both stuff i technically knew about but had gone completely blind to. the model doesn't have that. it reads the 300th file same as the first one.
only thing is that same freshness is why i don't trust it when it says something's fine. it catches what im numb to, but it'll also make stuff up with the same confidence. so i keep whatever it flags and ignore whatever it clears.
The signup-vs-webhook race on
usersis the useful finding. When you invent accounts to prove the fix, use reserved fiction phones (US 555-0100 to 555-0199, UK 020 7946 0xxx) and domains you control, so a leftover dump from the repro cannot publish a real person.