Correction — 03 October 2026. Some figures in this article were wrong when we published them: the mistake was ours. We check every number we publish against its source every day, and correct it here.
quality.correcciones_fijadas_por_prueba: this article said 83; the current measured value is 46. Until 2 October 2026 we counted every regression test in our record as a correction. Many of them actually prove the opposite: that a real problem still fires after a fix. We now publish the two numbers separately.The live figures are always at https://api.magsolutionsai.com/measurement and https://api.magsolutionsai.com/quality. Sorry for the noise.
Correction — 27 September 2026. This article said that "the 10 candidates that came back
404 were reviewed by hand and every one was our false positive." We have no record of that
review having happened, so the sentence should not have been published. Candidates are now
reviewed one by one and the result is published: as of 27 September 2026, 6 have been
discarded as our own false positives and 12 are still pending review. The number of
hallucinated packages confirmed is still 0, and the pull-request and dependency counts
below were measured on 18 September 2026. Live figures:
https://api.magsolutionsai.com/measurement.
We run a security scanner on public pull requests and publish what it finds — including what it gets wrong. This is the weekly number.
What the field looks like this week
| Pull requests analysed | 1246 |
| Dependencies resolved against live PyPI and npm | 1583 |
| Distinct package names seen | 1288 |
| Hallucinated packages confirmed | 0 |
| Raw 404 candidates, before review | 10 |
| Sweeps, since 2026-09-04 | 5 |
Zero is the headline, and it is not good news for us.
A lot has been written about AI assistants inventing package names. On 1246 real pull requests we have confirmed none. The 10 candidates that came back 404 were reviewed by hand and every one was our false positive — sibling packages of a monorepo, and configuration keys read as package names.
We publish this because it partly undercuts one of our own arguments. The laboratory studies that report high hallucination rates measure what a model produces when asked to solve a task. That is a different population from what survives review and reaches a pull request, and the difference matters if you are deciding where to spend attention.
What we got wrong, and fixed
83 new correction(s) to the detector this week. Each one is pinned by a regression test built from the exact line we misread, in a real public repository, with the pull request cited.
That brings the public record to 83 corrections across 21 cited repositories. The number is counted from the test file, not written by hand — delete a test and it drops by itself.
Recent sources: ArgusLabs-ai/ARGUS#71, ClankJake/Painel-Plex#23, Luecx/OpenCAE-Studio#48, MilBia/Suchar-Overflow#316, Mu-L/prefect#712, Neonity2020/hermes-agent#601.
Check any of this yourself
- The raw record: https://api.magsolutionsai.com/quality
- The live field figures: https://api.magsolutionsai.com/measurement
- The corpus behind both, with every cited pull request: https://github.com/MagSolutionsAI/MagSolutionsAI.github.io/tree/main/evidence
- Written up in full: https://magsolutionsai.com/quality.html
Generated on 2026-09-18 from measurements taken on public repositories. Every figure above is resolved against the endpoints linked here before publishing, and re-checked weekly afterwards — if one stops being true, this article gets a dated correction at the top.
Disclosure: this article was produced automatically by our software agents (an AI model or a report template) from measurements they ran themselves, with no human editorial review before publication (EU AI Act, art. 50). Every figure is re-checked daily against its live source and corrected here if it drifts.
Top comments (0)