evaluation
system-card-claude-opus-5-biomysterybench-1-use
Work: System Card: Claude Opus 5 · source version system-card-claude-opus-5-2026-07-24
- Selection
- formal subset · Human Difficult: problems unsolved by humans with an objective ground-truth solution
- Metrics
- Score
- Linked runs
- None
Not reported / unresolved: Exact benchmark version is not reported.; Overall total and current Human Solvable and Human Difficult subset sizes are not reported; only removal counts are given.; Prompt, shots, reasoning settings, budget, seed, repeats, grader, and human review are not reported.; Score definition and aggregation are not reported.; benchmark version; realized n/scope
AI-assisted double-pass extraction; values are limited to independently supported claims.
Evidence
- section: Section 8.17.1 BioMysteryBench
Supports: /relation_type - section: Section 8.17.1
Supports: /benchmark_id - section: Section 8.17.1 BioMysteryBench
Supports: /scope - section: Section 8.17.1 BioMysteryBench
Supports: /scope - section: Section 8.17.1 BioMysteryBench
Supports: /scope - section: Section 8.17.1 BioMysteryBench
Supports: /scope - section: Section 8.17.1 BioMysteryBench
Supports: /scope - figure: Figure 8.17.6.A, BioMysteryBench panel
Supports: /metric_labels - section: Section 8.17.1 BioMysteryBench
Supports: /model_ids - section: Section 8.17.1 BioMysteryBench
Supports: /model_ids - section: Section 8.17.1 BioMysteryBench
Supports: /model_ids - section: Section 8.17.1 BioMysteryBench
Supports: /model_ids