BioMysteryBench
An agentic bioinformatics benchmark of objective, expert-authored mysteries over anonymized real-world biological data, scored on final answers rather than prescribed analysis paths.
Scientific task · Scientific workflow
Complete a multi-step scientific analysis using data
端到端计算分析
end-to-end-computational-analysisCoverage
An agentic bioinformatics benchmark of objective, expert-authored mysteries over anonymized real-world biological data, scored on final answers rather than prescribed analysis paths.
A containerized benchmark of long-horizon bioinformatics analysis over real published notebooks and associated data, with open-answer and multiple-choice evaluation modes.
A cross-domain suite for discerning defensible analysis decisions and generating executable end-to-end analyses for open-ended scientific research questions; four of its twelve source questions are explicitly biological or ecological.
The BLADE track requiring a conceptual-variable specification, executable data-transformation function, and statistical-model function for each open-ended research question and dataset.
A 100-task agent benchmark of objectively gradable computational-biology problems requiring multi-step reasoning, bespoke code, tools, and real-world external resources.
A research-level agent benchmark of 129 synthetic, multistage computational-biology analyses that require iterative QC, statistical modeling, diagnostics, and decision-relevant judgment.
Each row keeps its original unit and basis. Rows with different units or overlapping mappings are never added.
| Benchmark | Mapped task | Coverage | Count | Version | Evidence |
|---|---|---|---|---|---|
| BioMysteryBench root: biomysterybench | End-to-end computational analysis official-taxonomy · high | explicitly-in-scope | 90 problems v11 mystery-bioinformatics problems after the June 2026 answer-key audit | v11 | biomysterybench-evidence-v11 |
| BixBench root: bixbench | End-to-end computational analysis official-taxonomy · high | explicitly-in-scope | 205 questions one question per row in the official v1.5 BixBench.jsonl | v1.5 | bixbench-evidence-v1-5-counts |
| BLADE root: blade | End-to-end computational analysis official-track · high | explicitly-in-scope | 12 problems paired real-world research questions and datasets used as the source units for BLADE | arXiv v3 | blade-evidence-current-counts |
| BLADE End-to-End Analysis Generation root: blade | End-to-end computational analysis official-track · high | explicitly-in-scope | 12 problems paired research questions and datasets requiring a complete generated analysis | arXiv v3 | blade-generation-evidence-counts |
| CompBioBench root: compbiobench | End-to-end computational analysis official-taxonomy · high | explicitly-in-scope | 100 tasks v1 independent computational-biology tasks | v1 | compbiobench-evidence-countscompbiobench-evidence-runner-license |
| GeneBench-Pro root: genebench-pro | End-to-end computational analysis official-taxonomy · high | explicitly-in-scope | 129 problems self-contained synthetic scientific-analysis problems (called evaluations in the paper abstract) | paper-v1 | genebench-pro-paper-evidence |
Runs are included only for benchmark records mapped here (and formal child tracks when a mapped suite is the root). A task mapping does not imply that every run isolates this task.