Relationship layer
Benchmark usage
This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.
evaluation
system-card-claude-opus-5-biomysterybench-1-use
Partialsubset
Benchmark: BioMysteryBench · version Not reported
- Selection
- formal subset · Human Difficult: problems unsolved by humans with an objective ground-truth solution
- Metrics
- Score
- Linked runs
- None
Not reported / unresolved: Exact benchmark version is not reported.; Overall total and current Human Solvable and Human Difficult subset sizes are not reported; only removal counts are given.; Prompt, shots, reasoning settings, budget, seed, repeats, grader, and human review are not reported.; Score definition and aggregation are not reported.; benchmark version; realized n/scope
AI-assisted double-pass extraction; values are limited to independently supported claims.
Evidence
- section: Section 8.17.1 BioMysteryBench
Supports: /relation_type - section: Section 8.17.1
Supports: /benchmark_id - section: Section 8.17.1 BioMysteryBench
Supports: /scope - section: Section 8.17.1 BioMysteryBench
Supports: /scope - section: Section 8.17.1 BioMysteryBench
Supports: /scope - section: Section 8.17.1 BioMysteryBench
Supports: /scope - section: Section 8.17.1 BioMysteryBench
Supports: /scope - figure: Figure 8.17.6.A, BioMysteryBench panel
Supports: /metric_labels - section: Section 8.17.1 BioMysteryBench
Supports: /model_ids - section: Section 8.17.1 BioMysteryBench
Supports: /model_ids - section: Section 8.17.1 BioMysteryBench
Supports: /model_ids - section: Section 8.17.1 BioMysteryBench
Supports: /model_ids
evaluation
system-card-claude-opus-5-proteingym-4-use
Partialunknown
Benchmark: ProteinGym · version Not reported
- Selection
- filtered · Subset of mutant protein sequences ranked against the wild type sequence
- Metrics
- rank correlation against real lab measurements
- Linked runs
- None
Not reported / unresolved: Exact ProteinGym version is not reported.; The Hard subset size and selection criteria are not reported.; Prompt, shots, reasoning settings, budget, seed, repeats, grader, and human review are not reported.; The rank-correlation variant and aggregation procedure are not reported.; benchmark version; realized n/scope
AI-assisted double-pass extraction; values are limited to independently supported claims.
Evidence
- section: Section 8.17.3 ProteinGym Hard
Supports: /relation_type - section: Section 8.17.3 ProteinGym Hard
Supports: /benchmark_id - section: Section 8.17.3 ProteinGym Hard
Supports: /scope - section: Section 8.17.3 ProteinGym Hard
Supports: /scope - section: Section 8.17.3 ProteinGym Hard
Supports: /scope - section: Section 8.17.3 ProteinGym Hard
Supports: /metric_labels - section: Section 8.17.3 ProteinGym Hard
Supports: /model_ids - section: Section 8.17.3 ProteinGym Hard
Supports: /model_ids - section: Section 8.17.3 ProteinGym Hard
Supports: /model_ids - section: Section 8.17.3 ProteinGym Hard
Supports: /model_ids
evaluation
system-card-claude-opus-5-scbench-3-use
Partialunknown · n=195
Benchmark: scBench · version Not reported
- Selection
- not reported
- Metrics
- Score
- Linked runs
- None
Not reported / unresolved: Exact benchmark version is not reported.; Per-workflow problem counts are not reported.; Prompt, shots, reasoning settings, budget, seed, repeats, grader, and human review are not reported.; Score definition and aggregation are not reported.; benchmark version
AI-assisted double-pass extraction; values are limited to independently supported claims.
Evidence
- section: Section 8.17.2 LatchBio Bioinformatics
Supports: /relation_type - section: Section 8.17.2 LatchBio Bioinformatics
Supports: /benchmark_id - section: Section 8.17.2 LatchBio Bioinformatics
Supports: /scope - section: Section 8.17.2 LatchBio Bioinformatics
Supports: /scope - figure: Figure 8.17.6.A, LatchBio Bioinformatics panel
Supports: /metric_labels - section: Section 8.17.2, SingleCellBench
Supports: /model_ids - section: Section 8.17.2, SingleCellBench
Supports: /model_ids - section: Section 8.17.2, SingleCellBench
Supports: /model_ids - section: Section 8.17.2, SingleCellBench
Supports: /model_ids
evaluation
system-card-claude-opus-5-spatialbench-2-use
Partialunknown · n=115
Benchmark: SpatialBench · version Not reported
- Selection
- not reported · Externally validated problems
- Metrics
- Score
- Linked runs
- None
Not reported / unresolved: The source does not identify an official benchmark version or artifact revision for the Verified qualifier.; Prompt, shots, reasoning settings, budget, seed, repeats, grader, and human review are not reported.; Score definition and aggregation are not reported.; benchmark version
AI-assisted double-pass extraction; values are limited to independently supported claims.
Evidence
- section: Section 8.17.2 LatchBio Bioinformatics
Supports: /relation_type - section: Section 8.17.2 LatchBio Bioinformatics
Supports: /benchmark_id - section: Section 8.17.2 LatchBio Bioinformatics
Supports: /scope - section: Section 8.17.2 LatchBio Bioinformatics
Supports: /scope - section: Section 8.17.2 LatchBio Bioinformatics
Supports: /scope - section: Section 8.17.2 LatchBio Bioinformatics
Supports: /scope - section: Section 8.17.2 LatchBio Bioinformatics
Supports: /metric_labels - section: Section 8.17.2, SpatialBench Verified
Supports: /model_ids - section: Section 8.17.2, SpatialBench Verified
Supports: /model_ids - section: Section 8.17.2, SpatialBench Verified
Supports: /model_ids - section: Section 8.17.2, SpatialBench Verified
Supports: /model_ids