It's 2 a.m. and I'm not hacking anything. I'm watching a list.
The list is made of package names that don't exist, recommended to thousands of developers by an assistant that sounds very sure of itself. All I have to do is register one of them and wait.
I don't need a zero-day. I don't need to phish anyone. The developer will install my code for me, because their AI told them to.
So I ran the experiment from the defender's side. How long would that list be?
Slopsquatting in 30 seconds
An AI coding tool suggests a package that doesn't exist. An attacker registers that exact name on npm or PyPI and puts something nasty inside. A developer copies the install command, and it works.
The term was coined by Seth Larson of the Python Software Foundation. It differs from typosquatting in one important way: no human makes a typo. The model makes the mistake, confidently.
Why I wanted my own numbers
Most posts on this topic are explainers or quote someone else's statistic. One recent preprint (not yet peer-reviewed) reported that five different LLMs invented the same 127 package names. If hallucinations repeat across models, they're predictable, and predictable means targetable.
It also means you can test it yourself. So I did.
The experiment
- Tools tested: Claude (Haiku 4.5), GitHub Copilot (auto/default model), ChatGPT / Codex (Codex default)
- Prompts: 30 everyday tasks, such as "parse a PDF", "validate an email", "rate-limit an API" (ecosystem: npm / JavaScript)
- Method: I extracted every package each tool suggested and checked whether it exists on the registry using a small script of my own that looked up each name on the npm registry
- Limit: 30 prompts is a small sample, so read these numbers as a signal, not a verdict.
- Safety rule: I never installed a package that failed the check. I only looked it up.
Results
| Tool | Fake packages found |
|---|---|
| Claude (Haiku 4.5) | 2 |
| GitHub Copilot (auto) | 1 |
| ChatGPT / Codex | 3 |
Across 30 prompts, the three tools made 6 fake package suggestions that don't exist on npm (counted per tool).
Overlap: 2 of those fake names showed up in more than one tool.
Suspicious but real: 2 packages existed but were very new or had almost no downloads. "It exists" is not the same as "it's safe." A package someone registered last week can be exactly what an attacker wants you to find.
I'm deliberately not publishing the fake names. A list of unregistered, AI-popular package names is a shopping list for attackers.
The part that surprised me
Two fake names appeared in more than one tool.
I expected noise. If each tool were simply guessing, two different products independently inventing the same nonexistent package should be rare. It happened twice in 30 prompts.
That's what turns a glitch into a pattern. A random mistake is hard to exploit. A repeatable one is a target.
So, hype or real?
Real, but not apocalyptic.
Six fake suggestions across 30 prompts is not a flood. Most of what these tools recommended was real, and my sample is small. But an attack like this only needs one hit:
- An AI invents a plausible name.
- Someone registers it before you do.
- A developer installs it without looking.
The two names that appeared in more than one tool are the ones I'd worry about, because repetition is what makes a hallucination worth squatting on. And the two suspicious-but-real packages show the second layer: even a successful lookup doesn't tell you who is behind the package.
A 5-minute defence checklist
- Never paste an AI-suggested install command blind. Look the package up first.
- Check age, downloads and repo link.
# npm: creation date and maintainers
npm view <package> time.created maintainers
# PyPI: metadata and release history
curl -s https://pypi.org/pypi/<package>/json | head -c 600
- Treat "new and unknown" as untrusted. A package created last week with a handful of downloads deserves a manual read.
- Commit your lockfile and review dependency diffs in PRs like you review code.
- Add a dependency scanner to CI so a bad name gets flagged before it ships.
A moment from my own test: one of the fake names looked so real that I didn't doubt it for a second. Nothing about it felt off. The registry lookup was the only thing that caught it. My instincts didn't.
Your turn
My process is the checklist above, and it only exists because one fake name in my test looked so real that my own instincts didn't flag it. The lookup did.
What's yours? How does your team check AI-suggested dependencies before they get installed: a process, a tool, or just trust?
Drop it in the comments. I'll collect the best answers into a follow-up.
Top comments (12)
Good that you were clear about the small sample and did not publish the fake names. Repeated hallucinations across tools is the finding that stands out. On your question, one cheap extra layer is a minimum release age setting in the package manager, since a freshly squatted package would be too new to install by default. It does not replace the manual age and maintainer check you describe, but it covers moments when someone skips it.
Thanks, Jessica, this is a genuinely useful addition. I didn't have the release-age gate in my checklist and it belongs there: npm has min-release-age from 11.10 (days, off by default), pnpm has minimumReleaseAge (minutes), and Yarn has npmMinimalAgeGate. I'll add it to the post with a credit.
One caveat I keep coming back to: the gate covers the fresh squat case, but a patient attacker who registers a commonly hallucinated name early and just waits would pass it once the package ages. So I see it as a speed bump for the race, not a fix for the pattern.
Are you running it per project or user-wide, and has it ever broken a CI build for you?
The two suspicious-but-real packages are the important limit of the registry lookup. I would separate dependency existence from dependency approval: resolve metadata without executing anything, then check the upstream repository and maintainer history before allowing a new package into the lockfile. For the next experiment, recording model version and total package suggestions would also make those six failures easier to compare across reruns without overstating what thirty prompts establish.
Thanks, Marcus. Existence vs approval is a better framing than the one in my post, and the two suspicious-but-real packages are exactly where a plain registry lookup stops being enough. Checking the upstream repo and maintainer history before anything reaches the lockfile makes sense.
You're also right about the gap: I didn't record total suggestions per tool, so I can't give a rate, and Copilot's auto mode means the model may have varied between prompts. For the next run I'll log the model version, total suggestions and package metadata per prompt, so reruns are actually comparable.
The overlap result is the part that matters here. One tool inventing a fake name is noise; two tools independently landing on the same name means it sits in some shared region of training-data plausibility, which is exactly what makes it worth registering.
One check we added on top of the basic lookup: compare
npm view <pkg> time.createdagainst roughly when the model suggested it. A name that resolves today but didn't exist a month ago is a worse signal than one that never resolved at all, because it may mean someone already squatted it. Publish date is nearly free to check and catches the case where you're the second victim rather than the first.Also agree on lockfile diffs in PR review. It's the cheapest place to catch this, since a brand-new dependency name is visually obvious in a diff in a way it never is inside an install command.
Thanks, this is a sharp addition. Comparing time.created against when the model suggested the name is the piece my lookup was missing: “resolves today, didn’t exist a month ago” says far more than exists. I didn’t timestamp when each name was suggested, so I can’t run that comparison on this run. Next time I will.
On the overlap: I’d call shared training-data plausibility a plausible hypothesis. My 30 prompts show the overlap but can’t show the cause.
On lockfile diffs, I agree it’s cheap and visible. Another commenter pointed out that by PR time an agent may already have run the install, so I’d treat the diff as the second gate, not the first. Do you run it as a CI check or by eye in review?
The most important part of this experiment is that package existence and package trust are two different verification steps. A registry lookup catches the obvious slopsquatting case, but a package resolving successfully doesn't mean it should be installed.
This is something we pay close attention to at IT Path Solutions when working with AI-assisted development: the model can propose a dependency, but the decision to introduce that dependency should happen outside the model. Registry existence, publisher identity, repository provenance, package age, lockfile changes, and ideally a CI policy should form the actual trust boundary.
The repeated hallucinations across tools are interesting too. Once multiple models converge on the same nonexistent name, the hallucination becomes more predictable and therefore potentially more exploitable. That makes maintaining an internal record of commonly suggested-but-invalid dependencies surprisingly useful.
I'd also add one gate before the install command: require the dependency to resolve to an approved registry artifact before the agent gets permission to execute the install at all. Detecting it in a later PR review is useful, but by then the agent has already crossed the supply-chain boundary.
Thanks, Mateo. Existence vs trust is the clearest way anyone has put it in this thread. One thing I'd add: a lookup only helps while the name is still unregistered. Once an attacker has squatted it, the lookup passes, which is exactly when the attack has worked. So the approval step has to carry the weight.
Your pre-install gate is the layer my checklist is missing. I'll add it, along with the idea of keeping an internal record of commonly suggested-but-invalid names (internal only, for the same reason I didn't publish mine).
Do you enforce the gate through agent permissions, or through an allowlisting proxy/private registry? Curious which one holds up better in practice.
The hallucination of package names is a massive security risk that often gets overlooked in the rush to adopt AI coding assistants. I've seen cases where a model suggests a perfectly plausible-sounding utility library that doesn't actually exist, and if a developer blindly runs a quick install, they are essentially inviting a typosquatting attack into their environment. It's not just about the code being wrong; it's about the supply chain vulnerability created by the model's desire to be helpful. I always make it a habit to verify any new dependency against a registry like npm or PyPI before adding it to my lockfile, even if the AI seems certain about it.
Thanks Anh, and thanks for answering the question I asked. Verifying against the registry before anything reaches the lockfile is the right habit, especially when the AI sounds certain.
Small note: this one is slopsquatting rather than typosquatting. With typosquatting the human mistypes a name; here the model invents the name, confidently.
One follow-up: besides checking that the package exists, do you also look at its age and maintainer? Two of the packages in my test existed but were suspiciously new.
The overlap number worries me most: 2 of the 6 invented names showing up in more than one tool means the hallucinations aren't random noise, they're predictable enough to pre-register. I've started gating my agent's install step on a registry check plus package age and downloads, since 'it exists' caught me out once with a week-old typo-twin. Did the 2 overlapping names repeat across separate runs, or only across tools?
Good question, and my data can't answer it: each prompt ran once per tool, so I only know the overlap across tools, not across reruns. That's a real gap. A name repeating across tools points to shared tendencies, while a name repeating across runs of the same tool would show stability, and stability is what makes pre-registering worth an attacker's time.
One clarification: the 6 is counted per tool, so the number of unique fake names is lower than 6. Two of them appeared in more than one tool.
For the next run I'll repeat each prompt several times per tool and track both. And thanks for the example of "it exists" catching you out with a week-old typo-twin. That's exactly the gap the age and downloads gate closes.