· via dev.to (home feed)
Hugging Face marks official benchmark datasets and ranks self-reported scores
Hugging Face now tags benchmark datasets as official and builds their leaderboards from self-reported eval files inside model repos, so a rank reflects the evaluation protocol as much as the model itself.

Hugging Face now flags selected benchmark datasets as official and assembles their leaderboards from evaluation files that model publishers upload to their own repositories, according to a dev.to post. Because every entry is self-reported, the headline number is only half the story: the sampling setup, voting scheme and thinking budget behind it determine whether two entries can be compared at all.
How the official boards are built
The mechanism has three parts, as the post describes it. A benchmark dataset carries a benchmark:official tag and ships an eval.yaml file defining its tasks. A model publisher then adds a file such as .eval_results/gpqa_diamond.yaml to the model repository, containing the dataset id, task id, value and date. The dataset page aggregates those files into a ranked leaderboard, and any board can be queried directly through Hugging Face's datasets API.
The catch is that the board displays the value but not necessarily how it was obtained. Protocol details, when they exist at all, sit in the notes field of the eval file and in the model card, and it is up to the reader to go looking for them.
The protocol behind five number-one scores
To illustrate, the post walks through Darwin-180B-RSI, a model hosted under the FINAL-Bench organisation, which it says ranks first on five of these official boards. The runs share a common configuration: a thinking budget of 131,072 tokens per sample, temperature 1.0, top_p 0.95, top_k 20, bf16 precision, and vLLM serving with tensor and expert parallelism.
The per-benchmark settings and results as reported:
| Benchmark | Samples | Reported score | Value |
|---|---|---|---|
| AIME 2026 | 16 | majority vote (mean 98.75) | 100.0 |
| HMMT Feb 2026 | 16 | majority vote (mean 96.59) | 100.0 |
| GPQA Diamond | up to 16 | majority vote | 94.44 |
| MMLU-Pro | 1 | single sample | 88.12 |
| MMMU-Pro (vision) | 3 | majority vote | 79.48 |
Three checks before comparing numbers
The post lists what to verify before treating two leaderboard entries as comparable:
- Samples and voting. A score built from a 16-way majority vote measures the whole sampling-and-voting setup, not a single model attempt, so it should not be lined up against single-sample numbers. Both figures need to be visible in the eval file.
- Thinking budget. Reasoning models lose points to truncation when the budget is small, and an answer cut short by a token limit is not the same as an answer the model got wrong.
- Contamination. Training data should be filtered against every test set a publisher reports on; the post credits VIDRAFT with an 8-gram overlap filter applied across the benchmarks it lists.
Reproducing the runs
All five benchmarks are public datasets, and the weights of Darwin-180B-RSI are openly available on Hugging Face. According to the post, serving the model in bf16 calls for eight B200-class GPUs, with four as the practical minimum, launched through vLLM with tensor parallelism, expert parallelism, a maximum context length of 135,168 tokens and trust-remote-code enabled.
Why it matters
An official tag makes benchmark results far easier to discover, but it does not make them uniform. Because every entry is self-reported, the sort order on a board silently mixes protocols: single-sample scores sit next to majority-vote scores, and a model evaluated with a generous thinking budget competes against one whose reasoning may have been truncated mid-stream. Verification work has shifted from a central operator to whoever reads the leaderboard. For teams choosing a model, the practical step is to open the eval file and model card before trusting a rank; for publishers, the format raises the expected standard of documenting exactly how each number was produced. Open weights, as in this case, at least keep independent reproduction possible when a claim needs to be checked.
- #hugging-face
- #benchmarks
- #llm
- #evaluation
- #open-weights