Protocol before rank

When do two scores belong in one comparison?

A benchmark name is a starting point, not proof of comparability. The ranking unit is a fingerprinted protocol + metric.

Four alignment steps

01

Fix data and split

Benchmark version, dataset version, split or challenge round, target subset, and label visibility must align.

02

Fix the metric definition

Macro/micro, mean/median, public/private, and first/best model are distinct. Scale conversions retain the original value and rule.

03

Fix training and inference

Zero-shot, probing, fine-tuning, external data, MSA/templates, ensembles, best-of-N, and inference budget split comparison groups.

04

Verify evidence and identity

Every numeric result needs a work, configuration, protocol, metric, and canonical evidence. Unlabelled chart estimates are excluded.

Four leader claims are not interchangeable

ClaimWhat it actually meansWhat it cannot imply
Official baselineOrganizer- or benchmark-designated reference.Not necessarily the weakest or earliest method.
Original-table bestBest value inside one source table.Cannot claim literature-wide SOTA.
Official-board leaderLeader in one official round and snapshot.Cannot be extrapolated across public/private boards or rounds.
Strict cross-work SOTABest across at least two independent works under a complete shared protocol.Unavailable when a critical field is unknown.

Ranking and gaps

  • Official ranks are preserved; recomputed ranks use competition ranking and keep every tie.
  • When the result inventory is incomplete, recomputed rank is explicitly an observed-subset rank.
  • No public exact value, restricted access, no standard numeric protocol, and unreconstructable protocol are separate states.
  • No cross-benchmark aggregate score or global model leaderboard is computed.

Snapshots and traceability

Every dated snapshot keeps a manifest, content hashes, chunk row counts, and literature cutoff. Dynamic leaderboard changes create a new snapshot instead of silently overwriting old scores.