Scientific task · Scientific workflow

End-to-end computational analysis

Complete a multi-step scientific analysis using data

端到端计算分析

Omics profileExperimental systemScientific workflow

Definition and search aliases

Permanent ID
end-to-end-computational-analysis
Aliases
agentic bioinformatics, computational investigation
Deprecated aliases
None
Hierarchy
Leaf task under Scientific workflow and systems analysis

Coverage

5 benchmark families cover this task

agentic-evalpartial

BioMysteryBench

An agentic bioinformatics benchmark of objective, expert-authored mysteries over anonymized real-world biological data, scored on final answers rather than prescribed analysis paths.

End-to-end computational analysis
agentic-evalpartial

BixBench

A containerized benchmark of long-horizon bioinformatics analysis over real published notebooks and associated data, with open-answer and multiple-choice evaluation modes.

End-to-end computational analysis
suitepartial

BLADE

A cross-domain suite for discerning defensible analysis decisions and generating executable end-to-end analyses for open-ended scientific research questions; four of its twelve source questions are explicitly biological or ecological.

End-to-end computational analysis
trackcomplete

BLADE End-to-End Analysis Generation

The BLADE track requiring a conceptual-variable specification, executable data-transformation function, and statistical-model function for each open-ended research question and dataset.

End-to-end computational analysis
agentic-evalpartial

CompBioBench

A 100-task agent benchmark of objectively gradable computational-biology problems requiring multi-step reasoning, bespoke code, tools, and real-world external resources.

End-to-end computational analysis
agentic-evalpartial

GeneBench-Pro

A research-level agent benchmark of 129 synthetic, multistage computational-biology analyses that require iterative QC, statistical modeling, diagnostics, and decision-relevant judgment.

End-to-end computational analysis

Evidence-backed count claims

Each row keeps its original unit and basis. Rows with different units or overlapping mappings are never added.

BenchmarkMapped taskCoverageCountVersionEvidence
BioMysteryBench
root: biomysterybench
End-to-end computational analysis
official-taxonomy · high
explicitly-in-scope90 problems
v11 mystery-bioinformatics problems after the June 2026 answer-key audit
v11biomysterybench-evidence-v11
BixBench
root: bixbench
End-to-end computational analysis
official-taxonomy · high
explicitly-in-scope205 questions
one question per row in the official v1.5 BixBench.jsonl
v1.5bixbench-evidence-v1-5-counts
BLADE
root: blade
End-to-end computational analysis
official-track · high
explicitly-in-scope12 problems
paired real-world research questions and datasets used as the source units for BLADE
arXiv v3blade-evidence-current-counts
BLADE End-to-End Analysis Generation
root: blade
End-to-end computational analysis
official-track · high
explicitly-in-scope12 problems
paired research questions and datasets requiring a complete generated analysis
arXiv v3blade-generation-evidence-counts
CompBioBench
root: compbiobench
End-to-end computational analysis
official-taxonomy · high
explicitly-in-scope100 tasks
v1 independent computational-biology tasks
v1compbiobench-evidence-counts
compbiobench-evidence-runner-license
GeneBench-Pro
root: genebench-pro
End-to-end computational analysis
official-taxonomy · high
explicitly-in-scope129 problems
self-contained synthetic scientific-analysis problems (called evaluations in the paper abstract)
paper-v1genebench-pro-paper-evidence

Official evaluations connected to these benchmarks

Runs are included only for benchmark records mapped here (and formal child tracks when a mapped suite is the root). A task mapping does not imply that every run isolates this task.

WorkProvider / classRelated runs
Evaluating Claude's bioinformatics research capabilities with BioMysteryBenchAnthropic
benchmark_creator
biomysterybench-official-run
biomysterybench-v8-human-difficult
biomysterybench-v8-human-solvable
BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational BiologyFutureHouse, ScienceMachine
benchmark_creator
bixbench-creator-paper
bixbench-paper-mcq-no-images
bixbench-paper-mcq-no-refusal
bixbench-paper-mcq-refusal
BixBench v1.5 dataset and evaluation releaseFutureHouse, ScienceMachine
benchmark_creator
bixbench-v1-5-agentic-mcq-no-refusal-images
bixbench-v1-5-agentic-mcq-refusal-images
bixbench-v1-5-agentic-mcq-refusal-no-images
bixbench-v1-5-agentic-open-images
bixbench-v1-5-zero-shot-mcq-no-refusal
bixbench-v1-5-zero-shot-mcq-refusal
bixbench-v1-5-zero-shot-open
BLADE: Benchmarking Language Model Agents for Data-Driven ScienceUniversity of Washington, UC Berkeley, New York University, Stanford University, University of British Columbia, Microsoft, George Washington University
benchmark_creator
blade-creator-decision-mcq
blade-creator-paper
blade-creator-react
Agentic systems are adept at solving well-scoped, verifiable problems in computational biologyGenentech, Roche
benchmark_creator
compbiobench-codex-hardest
compbiobench-creator-full
compbiobench-haiku-full
compbiobench-haiku-hardest
compbiobench-nonagentic-baselines
compbiobench-opus-full
compbiobench-opus-hardest
compbiobench-sonnet-full
compbiobench-sonnet-hardest
GeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational BiomedicineOpenAI
benchmark_creator
genebench-pro-claude-high
genebench-pro-claude-low
genebench-pro-claude-max
genebench-pro-claude-medium
genebench-pro-claude-xhigh
genebench-pro-official
genebench-pro-pro-mode
genebench-pro-reasoning-enabled
genebench-pro-standard-high
genebench-pro-standard-low
genebench-pro-standard-max
genebench-pro-standard-medium
genebench-pro-standard-none