Fix data and split
Benchmark version, dataset version, split or challenge round, target subset, and label visibility must align.
Protocol before rank
A benchmark name is a starting point, not proof of comparability. The ranking unit is a fingerprinted protocol + metric.
Four alignment steps
Benchmark version, dataset version, split or challenge round, target subset, and label visibility must align.
Macro/micro, mean/median, public/private, and first/best model are distinct. Scale conversions retain the original value and rule.
Zero-shot, probing, fine-tuning, external data, MSA/templates, ensembles, best-of-N, and inference budget split comparison groups.
Every numeric result needs a work, configuration, protocol, metric, and canonical evidence. Unlabelled chart estimates are excluded.
| Claim | What it actually means | What it cannot imply |
|---|---|---|
| Official baseline | Organizer- or benchmark-designated reference. | Not necessarily the weakest or earliest method. |
| Original-table best | Best value inside one source table. | Cannot claim literature-wide SOTA. |
| Official-board leader | Leader in one official round and snapshot. | Cannot be extrapolated across public/private boards or rounds. |
| Strict cross-work SOTA | Best across at least two independent works under a complete shared protocol. | Unavailable when a critical field is unknown. |
Every dated snapshot keeps a manifest, content hashes, chunk row counts, and literature cutoff. Dynamic leaderboard changes create a new snapshot instead of silently overwriting old scores.