BioMysteryBench
An agentic bioinformatics benchmark of objective, expert-authored mysteries over anonymized real-world biological data, scored on final answers rather than prescribed analysis paths.
Scientific task · Scientific workflow
Retrieve evidence
科学工作流与系统分析
scientific-workflow-systemsCoverage
An agentic bioinformatics benchmark of objective, expert-authored mysteries over anonymized real-world biological data, scored on final answers rather than prescribed analysis paths.
A containerized benchmark of long-horizon bioinformatics analysis over real published notebooks and associated data, with open-answer and multiple-choice evaluation modes.
A cross-domain suite for discerning defensible analysis decisions and generating executable end-to-end analyses for open-ended scientific research questions; four of its twelve source questions are explicitly biological or ecological.
The BLADE track requiring a conceptual-variable specification, executable data-transformation function, and statistical-model function for each open-ended research question and dataset.
A 100-task agent benchmark of objectively gradable computational-biology problems requiring multi-step reasoning, bespoke code, tools, and real-world external resources.
A research-level agent benchmark of 129 synthetic, multistage computational-biology analyses that require iterative QC, statistical modeling, diagnostics, and decision-relevant judgment.
A practical biology-research suite of 2,457 multiple-choice questions across eight broad categories and 31 versioned task files, with public and private contamination-monitoring splits.
Human-hard, multi-step multiple-choice scenarios involving plasmids, DNA fragments, enzymes, and molecular-cloning workflows.
Troubleshoots intentionally modified published biological protocols by selecting steps that would repair the stated outcome.
An agentic systems-biology suite in which language models iteratively perturb simulated SBML systems, analyze time-series observations in Python, and reconstruct hidden biological reactions.
The formally released SCIGYM track containing the 213 systems not included in the creator paper's model evaluation, with systems reaching up to 400 reactions.
The formally released and creator-evaluated SCIGYM track containing biological systems with fewer than ten reactions.
Database-retrieval category spanning 10 genomics, clinical, protein, regulatory, vaccine-response, and viral-PPI tasks.
Identifies genes associated with a phenotype in DisGeNET but not OMIM.
Retrieves human-gene cytogenetic locations from the stated Ensembl release.
Retrieves computationally predicted human miRNA targets from miRDB.
Retrieves genes in Mammalian Phenotype Tumor Ontology gene sets.
Retrieves membership in MSigDB C6 oncogenic-signature gene sets.
Retrieves promoter-region transcription-factor binding-site annotations from GTRD.
Uses a protein sequence and ClinVar lookup to identify benign or pathogenic variants.
Identifies ClinVar variant pathogenicity while reasoning across multiple protein sequences.
Retrieves membership in MSigDB vaccine-response gene sets.
Retrieves predicted human interaction partners of viral proteins from P-HIPSter.
The BLADE track for selecting the most or least justifiable conceptual-variable and data-transformation decisions for a research question and dataset.
Multiple-choice interpretation and multi-element reasoning over scientific figures shown without captions or paper context.
Literature-retrieval questions whose answers require findings in full research papers rather than titles or abstracts.
Retrieval and interpretation questions answerable from paper supplementary text or PDF tables.
Lookup, calculation, and reasoning questions over table images extracted from scientific papers.
Each row keeps its original unit and basis. Rows with different units or overlapping mappings are never added.
| Benchmark | Mapped task | Coverage | Count | Version | Evidence |
|---|---|---|---|---|---|
| BioMysteryBench root: biomysterybench | End-to-end computational analysis official-taxonomy · high | explicitly-in-scope | 90 problems v11 mystery-bioinformatics problems after the June 2026 answer-key audit | v11 | biomysterybench-evidence-v11 |
| BixBench root: bixbench | End-to-end computational analysis official-taxonomy · high | explicitly-in-scope | 205 questions one question per row in the official v1.5 BixBench.jsonl | v1.5 | bixbench-evidence-v1-5-counts |
| BLADE root: blade | End-to-end computational analysis official-track · high | explicitly-in-scope | 12 problems paired real-world research questions and datasets used as the source units for BLADE | arXiv v3 | blade-evidence-current-counts |
| BLADE End-to-End Analysis Generation root: blade | End-to-end computational analysis official-track · high | explicitly-in-scope | 12 problems paired research questions and datasets requiring a complete generated analysis | arXiv v3 | blade-generation-evidence-counts |
| CompBioBench root: compbiobench | End-to-end computational analysis official-taxonomy · high | explicitly-in-scope | 100 tasks v1 independent computational-biology tasks | v1 | compbiobench-evidence-countscompbiobench-evidence-runner-license |
| GeneBench-Pro root: genebench-pro | End-to-end computational analysis official-taxonomy · high | explicitly-in-scope | 129 problems self-contained synthetic scientific-analysis problems (called evaluations in the paper abstract) | paper-v1 | genebench-pro-paper-evidence |
| LAB-Bench root: lab-bench | Experiment and protocol planning official-track · high | explicitly-in-scope | 135 questions ProtocolQA questions across public and private splits. | repository-998a8e0 | lab-bench-evidence-paper |
| LAB-Bench CloningScenarios root: lab-bench | Experiment and protocol planning official-track · high | explicitly-in-scope | 41 questions questions across public and private splits | repository-998a8e0 | lab-bench-cloning-scenarios-evidence-paper |
| LAB-Bench ProtocolQA root: lab-bench | Experiment and protocol planning official-track · high | explicitly-in-scope | 135 questions questions across public and private splits | repository-998a8e0 | lab-bench-protocolqa-evidence-paper |
| SCIGYM root: scigym | Reaction-network reconstruction official-taxonomy · high | explicitly-in-scope | 350 systems distinct curated BioModels systems released as SBML benchmark instances | 2025 release | scigym-evidence-release-countsscigym-evidence-taxonomy |
| SCIGYM Large root: scigym | Reaction-network reconstruction official-track · high | explicitly-in-scope | 213 systems unique SBML systems in the official large Parquet split, containing the remaining systems with up to 400 reactions | 2025 release | scigym-large-evidence-count |
| SCIGYM Small root: scigym | Reaction-network reconstruction official-track · high | explicitly-in-scope | 137 systems unique SBML systems with fewer than 10 reactions in the official small Parquet split | 2025 release | scigym-small-evidence-count |
| LAB-Bench root: lab-bench | Scientific database retrieval official-track · high | explicitly-in-scope | 650 questions Questions across the ten formal DbQA child tasks. | repository-998a8e0 | lab-bench-evidence-repository |
| LAB-Bench DbQA root: lab-bench | Scientific database retrieval official-track · high | explicitly-in-scope | 650 questions questions across formal child tasks | repository-998a8e0 | lab-bench-dbqa-evidence-repository |
| LAB-Bench DbQA — Disease gene associations root: lab-bench | Scientific database retrieval official-track · high | explicitly-in-scope | 50 questions questions across public and private splits | repository-998a8e0 | lab-bench-dbqa-dga-evidence-paper |
| LAB-Bench DbQA — Gene location root: lab-bench | Scientific database retrieval official-track · high | explicitly-in-scope | 50 questions questions across public and private splits | repository-998a8e0 | lab-bench-dbqa-gene-location-evidence-paper |
| LAB-Bench DbQA — miRNA targets root: lab-bench | Scientific database retrieval official-track · high | explicitly-in-scope | 50 questions questions across public and private splits | repository-998a8e0 | lab-bench-dbqa-mirna-targets-evidence-paper |
| LAB-Bench DbQA — Mouse tumor gene sets root: lab-bench | Scientific database retrieval official-track · high | explicitly-in-scope | 100 questions questions across public and private splits | repository-998a8e0 | lab-bench-dbqa-mouse-tumor-gene-sets-evidence-paper |
| LAB-Bench DbQA — Oncogenic signatures root: lab-bench | Scientific database retrieval official-track · high | explicitly-in-scope | 50 questions questions across public and private splits | repository-998a8e0 | lab-bench-dbqa-oncogenic-signatures-evidence-paper |
| LAB-Bench DbQA — GTRD transcription-factor binding sites root: lab-bench | Scientific database retrieval official-track · high | explicitly-in-scope | 50 questions questions across public and private splits | repository-998a8e0 | lab-bench-dbqa-tfbs-gtrd-evidence-paper |
| LAB-Bench DbQA — Protein variant from sequence root: lab-bench | Scientific database retrieval official-track · high | explicitly-in-scope | 100 questions questions across public and private splits | repository-998a8e0 | lab-bench-dbqa-variant-from-sequence-evidence-paper |
| LAB-Bench DbQA — Protein variant with multiple sequences root: lab-bench | Scientific database retrieval official-track · high | explicitly-in-scope | 100 questions questions across public and private splits | repository-998a8e0 | lab-bench-dbqa-variant-multi-sequence-evidence-paper |
| LAB-Bench DbQA — Vaccine response gene sets root: lab-bench | Scientific database retrieval official-track · high | explicitly-in-scope | 50 questions questions across public and private splits | repository-998a8e0 | lab-bench-dbqa-vax-response-evidence-paper |
| LAB-Bench DbQA — Viral protein–protein interactions root: lab-bench | Scientific database retrieval official-track · high | explicitly-in-scope | 50 questions questions across public and private splits | repository-998a8e0 | lab-bench-dbqa-viral-ppi-evidence-paper |
| Anthropic Scientific Figure Interpretation Eval root: anthropic-key-life-sciences-evals | Scientific evidence interpretation official-track · high | explicitly-in-scope | Not reported private scientific figure interpretation tasks | reported-2026-01-11 as of 2026-01-11 | anthropic-scientific-figure-evidence |
| BLADE Decision-Discrimination MCQ root: blade | Scientific evidence interpretation official-track · high | explicitly-in-scope | 188 questions individual multiple-choice decision-discrimination questions | arXiv v3 | blade-mcq-evidence-counts |
| LAB-Bench root: lab-bench | Scientific evidence interpretation official-track · high | explicitly-in-scope | Not reported FigQA, LitQA2, SuppQA, and TableQA questions. | repository-998a8e0 | lab-bench-evidence-paper |
| LAB-Bench FigQA root: lab-bench | Scientific evidence interpretation official-track · high | explicitly-in-scope | 226 questions questions across public and private splits | repository-998a8e0 | lab-bench-figqa-evidence-paper |
| LAB-Bench LitQA2 root: lab-bench | Scientific evidence interpretation official-track · high | explicitly-in-scope | 248 questions questions across public and private splits | repository-998a8e0 | lab-bench-litqa2-evidence-paper |
| LAB-Bench SuppQA root: lab-bench | Scientific evidence interpretation official-track · high | explicitly-in-scope | 102 questions questions across public and private splits | repository-998a8e0 | lab-bench-suppqa-evidence-paper |
| LAB-Bench TableQA root: lab-bench | Scientific evidence interpretation official-track · high | explicitly-in-scope | 305 questions questions across public and private splits | repository-998a8e0 | lab-bench-tableqa-evidence-paper |
| SCIGYM root: scigym | Simulation-based experiment official-taxonomy · high | explicitly-in-scope | 350 systems distinct curated BioModels systems released as SBML benchmark instances | 2025 release | scigym-evidence-release-countsscigym-evidence-taxonomy |
| SCIGYM Large root: scigym | Simulation-based experiment official-track · high | explicitly-in-scope | 213 systems unique SBML systems in the official large Parquet split, containing the remaining systems with up to 400 reactions | 2025 release | scigym-large-evidence-count |
| SCIGYM Small root: scigym | Simulation-based experiment official-track · high | explicitly-in-scope | 137 systems unique SBML systems with fewer than 10 reactions in the official small Parquet split | 2025 release | scigym-small-evidence-count |
Runs are included only for benchmark records mapped here (and formal child tracks when a mapped suite is the root). A task mapping does not imply that every run isolates this task.