Hugging Face marks a growing set of datasets as official benchmarks. Each one gets a leaderboard on its dataset page, filled automatically from the .eval_results files that model repositories publish. There are now 48 of them, from GPQA Diamond and MMLU-Pro to SWE-bench, Terminal-Bench, AIME 2026 and the Open ASR Leaderboard.
What has been missing is a way to see them together. Each leaderboard lives on its own page, so questions like "which fields do these benchmarks cover?" or "where do model makers actually submit?" are hard to answer.
A new Space, Hugging Face Official Benchmarks: Map and Participation, pulls every official leaderboard into one live page. It reads the public Hugging Face API in the browser on each visit, so the numbers below are a snapshot from October 4, 2026, and the page itself always shows the current state.
Here is what the map shows.
1. The benchmarks lean toward agents
Grouping the 48 benchmarks by field:
| Field | Official benchmarks |
|---|---|
| Agents and terminal tasks | 17 |
| Science and knowledge | 6 |
| Documents and OCR | 5 |
| Coding | 4 |
| Vision and video | 4 |
| Law and finance | 4 |
| Math | 3 |
| Speech | 2 |
| Retrieval | 2 |
| Robotics | 1 |
More than a third of the official set measures agents: terminal work, tool use, long-horizon tasks, browsing and multi-step workflows. That matches where the industry conversation has moved over the past year.
2. Submissions go somewhere else
The picture changes when you count leaderboard entries instead of benchmarks. Of roughly 1,020 entries:
| Field | Entries |
|---|---|
| Science and knowledge | 348 |
| Agents and terminal tasks | 197 |
| Coding | 158 |
| Documents and OCR | 89 |
| Speech | 72 |
| Math | 62 |
| Vision and video | 47 |
| Law and finance | 39 |
| Retrieval | 4 |
| Robotics | 4 |
Agents have the most benchmarks, but science and knowledge get the most submissions. Part of the reason is practical: GPQA and MMLU-Pro are cheap to run and widely reported, while agent benchmarks need harnesses, sandboxes and long runs. Many new agent benchmarks still have only a handful of entries.
3. Participation is very uneven
- The median official benchmark has about 12 leaderboard entries.
- 14 of them have five entries or fewer.
- The five busiest boards (MMLU-Pro about 141 entries, GPQA about 111, HLE about 89, SWE-bench Verified about 67, Open ASR about 62) hold 46% of all entries.
So a small group of benchmarks carries most of the comparison, and the long tail is thin. A #1 position on a board with 140 entries and one on a board with 3 entries are very different things, which is worth keeping in mind whenever a leaderboard result is quoted.
4. Who submits
About 410 models from 95 organizations appear across the official leaderboards. The most active publishers by number of entries:
| Organization | Entries |
|---|---|
| Qwen | 136 |
| Z.ai (zai-org) | 84 |
| DeepSeek | 65 |
| Moonshot AI | 65 |
| NVIDIA | 63 |
| Ornith | 48 |
| OpenAI | 47 |
| MiniMax | 36 |
Grouped by the country of the publishing organization, about 55% of entries come from China, 23% from the USA, 4% from Korea and 2% from France, with 13% from organizations without a clear public headquarters. These country labels are assigned by hand in the Space and are open for correction.
5. What else is on the page
- #1 spots: for every benchmark, which model is ranked first and how many first places each organization holds, plus a heatmap of who leads which field.
- Gaps: how far #1 is ahead of #2 on each board, in percent.
- Boards entered vs. boards led: whether an organization wins by entering many boards or by winning most of the few it enters.
- Every benchmark, sorted by number of entries, with #1 and #2.
How to read it
Two cautions apply to any view built on these leaderboards.
First, entries are published by model authors in their own repositories, and settings differ: sampling, majority voting, thinking budget and prompt format all vary between submissions. A leaderboard ranks the numbers authors report; it does not run the models itself.
Second, an organization is the namespace of the model repository. A fine-tune published under a different namespace counts for that namespace, not for the base model's creator.
With those in mind, the map is useful for a different kind of question than a single leaderboard answers: not "who is best at X", but "what is being measured, how crowded each measurement is, and who is showing up".
The Space: https://huggingface.co/spaces/quantid/huggingface-official-benchmark-leaderboards
Top comments (0)