Correction — 03 October 2026. Some figures in this article were wrong when we published them: the mistake was ours. We check every number we publish against its source every day, and correct it here.
quality.correcciones_fijadas_por_prueba: this article said 83; the current measured value is 46. Until 2 October 2026 we counted every regression test in our record as a correction. Many of them actually prove the opposite: that a real problem still fires after a fix. We now publish the two numbers separately.The live figures are always at https://api.magsolutionsai.com/measurement and https://api.magsolutionsai.com/quality. Sorry for the noise.
We run a security scanner on public pull requests and publish what it finds — including what it gets wrong. This is the weekly number.
What the field looks like this week
| Pull requests analysed | 3734 |
| Dependencies resolved against live PyPI and npm | 3500 |
| Distinct package names seen | 3691 |
| Hallucinated packages confirmed | 0 |
| Candidates discarded on review | 6 |
| Candidates still pending review | 12 |
| Sweeps, since 2026-09-04 | 15 |
A 404 means a name is not published. It does not mean anyone invented it — an unpublished sibling inside a monorepo returns exactly the same 404, and no HTTP request can tell them apart. So a name is counted only once a person has confirmed it, and the discarded and pending counts are published beside the total.
12 candidate(s) have not been looked at yet and are not in the count above. Uncounted is the default: we would rather say nothing than say something we have not checked.
Zero confirmed is the headline, and it is not good news for us.
A lot has been written about AI assistants inventing package names. On 3734 real pull requests we have confirmed none, and discarded 6 candidate(s) on review — sibling packages of a monorepo, and configuration keys read as package names.
We publish this because it partly undercuts one of our own arguments. The laboratory studies that report high hallucination rates measure what a model produces when asked to solve a task. That is a different population from what survives review and reaches a pull request, and the difference matters if you are deciding where to spend attention.
Check any of this yourself
- The raw record: https://api.magsolutionsai.com/quality
- The live field figures: https://api.magsolutionsai.com/measurement
- The corpus behind both, with every cited pull request: https://github.com/MagSolutionsAI/MagSolutionsAI.github.io/tree/main/evidence
- Written up in full: https://magsolutionsai.com/quality.html
Generated on 2026-09-28 from measurements taken on public repositories. Every figure above is resolved against the endpoints linked here before publishing, and re-checked weekly afterwards — if one stops being true, this article gets a dated correction at the top.
Disclosure: this article was produced automatically by our software agents (an AI model or a report template) from measurements they ran themselves, with no human editorial review before publication (EU AI Act, art. 50). Every figure is re-checked daily against its live source and corrected here if it drifts.
Top comments (0)