DEV Community

Cover image for 1 in 5 Packages Your AI Suggests Don't Exist. Attackers Know Which Ones.

1 in 5 Packages Your AI Suggests Don't Exist. Attackers Know Which Ones.

James Anderson on September 30, 2026

You ask your AI assistant how to do something. It gives you clean, confident code, with an install line at the top: pip install aws-helper-sdk ...
Collapse
 
xxxn3m3s1sxxx profile image
xxxn3m3s1sxxx • • Edited

The "pin and hash" advice has a hole in it that only shows up once agents are the ones running the install: the lockfile is written by the same process that hallucinated the name. A pin you generated yourself is a self-attestation, not a check — the agent adds the package, resolves it, commits the lockfile, and the hash now faithfully records a bad decision. Nothing in the diff looks wrong either, because a lockfile entry is a wall of checksums no reviewer actually reads.

So the useful property isn't "is it pinned," it's "who was allowed to change the pin." What's worked for us is treating the dependency set as a reviewed artifact rather than a build byproduct: adding a name is a proposed change that a different process validates against the registry before it lands, and the agent that suggested the name can't be the one that approves its own resolution. Age-gating and allowlists both slot in naturally there, because that's the chokepoint where the request exists before it's already true.

The other half is that verification has to re-run outside the agent's write path — CI reading the lockfile from the repo is fine, CI trusting whatever the agent just reported is the same trust problem one layer up.

Collapse
 
slabb profile image
Sam LABBE •

The two-writer rule is the fix; the part that makes it survive an incident is turning the writers' outputs into evidence, not just separating them. Proposal, approval, and the CI re-resolution as three events from three writers, correlated by the resolution they refer to — when they disagree after the fact, you have a sequence to read, not an opinion to negotiate. Because the recursion in your last line keeps going otherwise: the CI that re-runs the check is itself a process whose report can be wrong, one layer up each time. What ends the regress is sealing each claim at write time — the chain doesn't vouch for truth, it vouches for who claimed what and when, which is what makes a silent bypass legible as "approval with no matching proposal."

Collapse
 
james_anderson_h profile image
James Anderson •

Turning outputs into evidence is what makes the two-writer rule survive the incident — proposal, approval, CI re-resolution as three correlated events gives you a sequence to read, not an opinion to negotiate. And you've ended the regress: verification can't be what you trust, because the verifier's report can be wrong one layer up, forever. Sealing each claim at write time vouches for who claimed what and when, not truth — so a silent bypass shows up as a hole in the sequence, not a judgment call.

Collapse
 
james_anderson_h profile image
James Anderson •

The sharpest hole anyone's found: when the agent runs the install, the lockfile is written by the same process that hallucinated the name, so the pin becomes self-attestation — the hash faithfully records a bad decision, and nobody reads a wall of checksums. Self-grading, one layer down. The fix is "who's allowed to change the pin": the dependency set as a reviewed artifact, validated by a different process before it lands. Verification must re-run outside the agent's write path.

Collapse
 
normalnorma profile image
Norma •

Great breakdown. The shift you describe, from the vulnerability being your carelessness to being your trust, is what makes slopsquatting different from typosquatting. Similarity-based detection is built around human typos, so it has nothing to say about a plausible-sounding name like aws-helper-sdk that a model invented.

The predictability finding is the part that turns this from an annoyance into an attack. If 43% of hallucinated names recur on every rerun, an attacker can harvest them systematically and register them in advance. The cross-model result, where five frontier models independently invented the same 127 names, makes that worse, because switching assistants doesn't protect you.

Your point about agents is the one I'd stress most. A human at least has a moment where they might glance at the package name. An autonomous agent that runs its own installs removes that moment entirely, so gating installs behind verification or approval matters more than any of the other controls.

A few additions to the defense list:

  • Check package age and publish date, not just existence. A package that was first published a few weeks ago and matches a name a model just suggested is a strong signal. Some tooling can flag "newly registered and matches an LLM suggestion" automatically.
  • Use install-time restrictions where the ecosystem allows it. Disabling install scripts by default (e.g. npm install --ignore-scripts) limits what a malicious package can do on arrival, even if it slips through.
  • Prefer name verification against the official docs. If an assistant suggests a package for a well-known library, confirm the name on that library's own documentation or repo rather than asking the model again, since it will often repeat the same hallucination.
  • Reserve your own likely names. For maintainers, registering obvious variants of your project's name (yourproject-sdk, yourproject-cli) closes off the cheapest squatting targets.

One question: do you know whether the recurrence rate differs between prompts that mention a specific framework and generic "how do I do X" prompts? My guess is that the vague, task-shaped prompts produce the more reliably exploitable names, since there's no real package to anchor on, but I haven't seen data either way.

The ten-second rule at the end, treating a package name from an AI as a claim to verify rather than a fact to run, is easy to remember and cheap to adopt. Thanks for the clear write-up and the sourcing.

Collapse
 
james_anderson_h profile image
James Anderson •

Those four additions strengthen the defense list materially, and --ignore-scripts is the one I most regret leaving out — it bounds what a malicious package can do on arrival, which is defense-in-depth even for the squat that slips the net. Reserving your own likely names is the cheap maintainer-side move nobody thinks of until it's too late. On your question: I don't have hard data splitting recurrence by prompt type, but your hypothesis matches the mechanism exactly — vague task-shaped prompts have no real package to anchor on, so the model fills the gap with its most probable invention, which is precisely the stable, recurring, registerable kind. Framework-specific prompts anchor on something real and fail toward the correct name more often. If that holds, the most exploitable hallucinations come from exactly the beginner-style "how do I do X" queries — which is the worst possible population to be exposed. Worth someone measuring directly.

Collapse
 
pepapepa profile image
pepapepa •

I don't like something around this user. trouble. bad mojo vibes.

Collapse
 
slabb profile image
Sam LABBE •

This is the "see you there" arriving, then — the buildable half, wearing a supply-chain badge.

Honest answer first: yes, more than once — and almost always in a throwaway environment, which is exactly the rationalization this attack is built around. "The build worked, I moved on" is the whole story.

The thing I'd push one step further: the predictability you flag cuts both ways. If 43% of hallucinated names recur on every rerun, then your model's hallucination set is finite and enumerable. Run your real prompt corpus, collect every package name it invents, and that list becomes regression fixtures and a seed denylist for your registry gate. Attackers harvest your model's hallucinations from the outside; the cheap defense is to harvest them from the inside, from your own traffic, first.

On "gate its installs" — a gate that leaves no record is a habit, not a control. After an incident, "we had a verification step" is a policy statement; what an auditor (or an insurer) needs is the sequence itself: which model emitted the name (a claim), what the registry check returned (a verification), what approved or refused the install (a decision) — three separate events, reconcilable after the fact, the way you'd reconcile a payment decision against a provider response. That's the pattern behind the flight-recorder journal I keep mentioning: an agent that can install is an agent that must journal.

And your docs point might be the sleeper: agents increasingly write the docs too, so the verification has to attach to the artifact and re-run in CI — not just to the moment someone pastes the line.

Nothing caught in the wild on my side yet. But the early-warning metric is now obvious: the fake name that comes back on every rerun is the one to watch.

Collapse
 
james_anderson_h profile image
James Anderson •

The inversion is the sharpest thing added: if 43% of hallucinations recur, the model's hallucination set is finite and enumerable — so run your own prompt corpus, harvest what it invents, and that list becomes regression fixtures and a seed denylist. Attackers harvest your model's hallucinations from outside; the cheap defense is beating them to it from the inside. That flips a scary stat into a build step.

And you're right that a gate with no record is a habit, not a control — three reconcilable events (claim, verification, decision) is the difference between "we had a step" and evidence. An agent that can install is an agent that must journal. See you there, indeed.

Collapse
 
slabb profile image
Sam LABBE •

Twice now, then — "see you there" has closed both of these. So, the address: the journal's reconciliation layer, where claim, verification and decision become typed, replayable events, is sitting open as an issue that needs skeptics more than applause. You know where it lives. Bring the attempted breaks.

Thread Thread
 
james_anderson_h profile image
James Anderson •

Noted the address — I'll come with breaks, not applause.

Collapse
 
anciwasim profile image
Wasim Sheikh •

Long-lived trust is how a generated package name turns into a supply-chain problem after the human has left the chat. We tie dependency approval to the task, verify the package and owner, and hard-stop when that envelope expires. “The model suggested it” is not a permission model.

Collapse
 
james_anderson_h profile image
James Anderson •

"The model suggested it is not a permission model" — that's the whole fix in a sentence. Tying approval to the task and hard-stopping when the envelope expires kills the long-lived-trust gap, because the danger isn't the suggestion, it's the standing permission that outlives the human who should've checked it.

Collapse
 
build996 profile image
build996 •

Most defenses in this thread check the name. The attack also has a tell that doesn't depend on the name at all: when the package came into existence. A squatted hallucination is younger than the model habit that invents it, so a CI gate that holds any new dependency whose first release is under 90 days old covers the long tail raised above without anyone harvesting names first. npm exposes time.created and PyPI's JSON API lists upload times per release, so it's a few lines. It stops working once a squatted package has aged past the threshold, which is exactly the 233-downloads-a-week case, so the allow-list still has to carry that one.

Collapse
 
james_anderson_h profile image
James Anderson •

This is the sharpest defense in the thread, because it sidesteps the name entirely — you don't need to know which hallucination, you just exploit the one invariant the attacker can't fake: a squatted package is always younger than the model habit that invents it. Gating any new dependency whose first release is under ~90 days old catches the whole long tail without harvesting a single name first, and it's a few lines against npm's time.created and PyPI's upload timestamps. The honest limit you flag is the right one: it fails exactly on the aged squat (the 233-a-week case), so age-gating and allow-listing aren't competitors — age handles the fresh long tail cheaply, the allow-list carries the patient attacker who waited out the window. Layer the two and you've covered both the lazy and the disciplined adversary. Stealing this.

Collapse
 
aidiveyt profile image
AI Dive •

The gate at the end is cheap to build. My 42 line hook on the tool call path, written to redact secrets, caught both Read and Bash output at a 27.9 ms round trip. The same seam takes an install command: check registry age and downloads, deny anything published last week.

Collapse
 
james_anderson_h profile image
James Anderson •

That's the whole argument made concrete — the expensive-sounding defense is a 42-line hook on the tool-call path, and the fact that it already caught both Read and Bash output at ~28ms proves the seam is general, not purpose-built. The install-gate is just another policy on the same chokepoint: intercept the command, check registry age and download count, deny anything a week old. The age-based rejection is especially nice there because it catches the fresh squat without needing to know the name in advance — the attacker can't make a just-registered package look old. The lesson underneath: once you have one enforcement point the model can't talk its way past, every new threat becomes a rule you add there rather than a system you rebuild. Cheap seam, reusable boundary. That's the part people underestimate because it isn't glamorous.

Collapse
 
indiainfranotes profile image
IndiaInfraNotes •

The package manager is a trust boundary, not a typing detail. A practical guardrail is to resolve the exact package and version through the registry API, fail closed on a miss, then pin the lockfile and review install scripts. That catches both hallucinated names and last-minute takeover attempts.

Collapse
 
james_anderson_h profile image
James Anderson •

Exactly — "trust boundary, not a typing detail" is the reframe that should reorganize how people think about this. The install line isn't a convenience to breeze past; it's the exact seam where untrusted names cross into your running system.

Your guardrail is strong because it fails closed: resolve the exact name and version through the registry first, and a hallucinated package dies at the boundary before anything executes — prevention, not detection. Pin the lockfile to kill the last-minute takeover; review scripts to bound what survives.

Best part: it's all deterministic. No classifier, no arms race. Exists and resolves — yes or no.

Collapse
 
docify profile image
Docify •

Hey, I'm building a small open-source CLI that analyzes a codebase and generates architecture/structure documentation. I'm looking for a few developers willing to run it against a real project and tell me where it gets things wrong.
You don't need to upload your code anywhere just run:
npx @autodocify/autodocs analyze .
Requires Node 20+.
If you try it, I'd especially like to know what it missed or misunderstood.

Collapse
 
kartik-nvjk profile image
Kartik N V J K •

The 19.7% hallucination rate across 16 models is bad enough, but the 43% that recur on every re-run is the detail that makes this weaponizable, an attacker only needs the name to be predictable. Registering 205,474 plausible package names isn't hard once the model keeps suggesting the same fakes. Are you screening install-time against a known-good allowlist, or catching these at PR review?

Collapse
 
james_anderson_h profile image
James Anderson •

You've put your finger on exactly why the 43% matters more than the 19.7% — randomness would protect you; predictability is what makes it a registerable target. On your question: allowlist at install-time is the stronger control, because PR review relies on a human recognizing an unfamiliar name, and the whole problem is that these sound legitimate. Catch it at the gate, not the eyeball.

Collapse
 
evanbright profile image
Evan Bright •

The part about treating AI-generated package names as claims rather than facts really stands out. I think dependency verification should become a normal part of any AI-assisted coding workflow, especially when agents can install packages automatically.

Collapse
 
james_anderson_h profile image
James Anderson •

Exactly — "verify before install" has to become reflex, not an afterthought, the moment agents can install on their own.

Collapse
 
julianneagu profile image
Julian Neagu •

The agent install step is the bit I’d worry about most. I’ve built tools where adding a package is basically one command, so a registry allowlist feels worth the extra friction.

Collapse
 
james_anderson_h profile image
James Anderson •

Exactly — the moment adding a package is one frictionless command, the friction is the safety, and a registry allowlist is the cheapest place to put it back. An agent that can install unvetted packages at machine speed is the highest-risk version of this whole problem, so trading a little convenience for "only approved names resolve" is a good deal.

Collapse
 
Sloan, the sloth mascot
Comment deleted
Collapse
 
slabb profile image
Sam LABBE •

The tail is real, but the harvest was never meant to come from the study's 205k — it comes from your own corpus, on your own prompts, so the set that matters is whatever your stack actually re-invents. And the same corpus is the drift detector: re-run it on every model upgrade, diff the phantom set, and the gate learns the new names before production does. Allow-list stays the strong control; the harvest is what keeps it current.

Collapse
 
james_anderson_h profile image
James Anderson •

Right — the 205k is the global tail; your corpus only has to cover what your stack re-invents, which is a far smaller, finite set. And diffing the phantom set on every model upgrade is the part I underweighted — the harvest isn't a one-time seed, it's a drift detector that catches the new hallucinations before prod does. Allow-list bounds the damage; the harvest keeps it from going stale the moment you bump a model version. The upgrade is the attack window, so that's exactly when to re-run it.

Thread Thread
 
slabb profile image
Sam LABBE •

The upgrade-window point is the one to operationalize: run the harvest against the candidate model as part of the eval pass, before the bump ships. That way the diff lands in the release review, while the delta is still cheap to act on — and the allow-list gets versioned with the model it gates.

Thread Thread
 
james_anderson_h profile image
James Anderson •

Running the harvest in the eval pass against the candidate model — before the bump ships — is exactly the right place to put it, because it moves the drift from a post-deploy surprise to a release-review artifact while the delta's still cheap. The model upgrade is the attack window, so gating it there means the new hallucinations surface in the same review that approves the model, not in production a week later. And versioning the allow-list with the model it gates is the detail most people miss: the phantom set isn't a global constant, it's a property of a specific model version, so pinning them together means a rollback takes the right denylist with it. The harvest becomes a build step, the diff becomes a review gate, and the allow-list stops drifting the moment you bump the version.

Thread Thread
 
slabb profile image
Sam LABBE •

Rollback taking the right denylist with it is the phrasing that goes in the docs — the phantom set as a property of a model version, not a global constant. Door's open; bring the breaks.

Collapse
 
jj1423 profile image
jj1423 •

The edit-distance point is the one I'd underline. "Only ~13% were simple
typos" explains why this slipped past tooling that was working correctly —
similarity detection was never wrong, it was answering a different question.

We see the same thing from the other end. We run a fixed set of coding tasks
through a few models and look up every name that comes back, and the invented
names don't cluster around real package names — they cluster around the task.
Ask for CycloneDX parsing in Python and you get cyclonedx-pythonlib,
cyclonedx-python3, cyclonedx-xml-python. Nothing a Levenshtein check would
pair with anything, and nothing a reviewer would stop at.

Which is your carelessness-to-trust shift in one line: a typo is a name that
looks slightly wrong, a hallucination is a name that looks exactly right.

Names and the per-task breakdown: github.com/0pstech/slopsquatting-dataset
(CC BY 4.0)