Scientific task · Scientific workflow

Scientific workflow and systems analysis

Retrieve evidence

科学工作流与系统分析

Experimental systemOmics profileScientific workflow

Definition and search aliases

Permanent ID
scientific-workflow-systems
Aliases
scientific analysis workflow
Deprecated aliases
None
Hierarchy
Parent task; benchmark coverage below includes its leaf tasks.

Coverage

7 benchmark families cover this task

agentic-evalpartial

BioMysteryBench

An agentic bioinformatics benchmark of objective, expert-authored mysteries over anonymized real-world biological data, scored on final answers rather than prescribed analysis paths.

End-to-end computational analysis
agentic-evalpartial

BixBench

A containerized benchmark of long-horizon bioinformatics analysis over real published notebooks and associated data, with open-answer and multiple-choice evaluation modes.

End-to-end computational analysis
suitepartial

BLADE

A cross-domain suite for discerning defensible analysis decisions and generating executable end-to-end analyses for open-ended scientific research questions; four of its twelve source questions are explicitly biological or ecological.

End-to-end computational analysis
trackcomplete

BLADE End-to-End Analysis Generation

The BLADE track requiring a conceptual-variable specification, executable data-transformation function, and statistical-model function for each open-ended research question and dataset.

End-to-end computational analysis
agentic-evalpartial

CompBioBench

A 100-task agent benchmark of objectively gradable computational-biology problems requiring multi-step reasoning, bespoke code, tools, and real-world external resources.

End-to-end computational analysis
agentic-evalpartial

GeneBench-Pro

A research-level agent benchmark of 129 synthetic, multistage computational-biology analyses that require iterative QC, statistical modeling, diagnostics, and decision-relevant judgment.

End-to-end computational analysis
suitepartial

LAB-Bench

A practical biology-research suite of 2,457 multiple-choice questions across eight broad categories and 31 versioned task files, with public and private contamination-monitoring splits.

Scientific database retrievalScientific evidence interpretationExperiment and protocol planning
trackcomplete

LAB-Bench CloningScenarios

Human-hard, multi-step multiple-choice scenarios involving plasmids, DNA fragments, enzymes, and molecular-cloning workflows.

Experiment and protocol planning
trackcomplete

LAB-Bench ProtocolQA

Troubleshoots intentionally modified published biological protocols by selecting steps that would repair the stated outcome.

Experiment and protocol planning
suitecomplete

SCIGYM

An agentic systems-biology suite in which language models iteratively perturb simulated SBML systems, analyze time-series observations in Python, and reconstruct hidden biological reactions.

Reaction-network reconstructionSimulation-based experiment
trackcomplete

SCIGYM Large

The formally released SCIGYM track containing the 213 systems not included in the creator paper's model evaluation, with systems reaching up to 400 reactions.

Reaction-network reconstructionSimulation-based experiment
trackcomplete

SCIGYM Small

The formally released and creator-evaluated SCIGYM track containing biological systems with fewer than ten reactions.

Reaction-network reconstructionSimulation-based experiment
trackcomplete

LAB-Bench DbQA

Database-retrieval category spanning 10 genomics, clinical, protein, regulatory, vaccine-response, and viral-PPI tasks.

Scientific database retrieval
trackcomplete

BLADE Decision-Discrimination MCQ

The BLADE track for selecting the most or least justifiable conceptual-variable and data-transformation decisions for a research question and dataset.

Scientific evidence interpretation
trackcomplete

LAB-Bench FigQA

Multiple-choice interpretation and multi-element reasoning over scientific figures shown without captions or paper context.

Scientific evidence interpretation
trackcomplete

LAB-Bench LitQA2

Literature-retrieval questions whose answers require findings in full research papers rather than titles or abstracts.

Scientific evidence interpretation
trackcomplete

LAB-Bench SuppQA

Retrieval and interpretation questions answerable from paper supplementary text or PDF tables.

Scientific evidence interpretation
trackcomplete

LAB-Bench TableQA

Lookup, calculation, and reasoning questions over table images extracted from scientific papers.

Scientific evidence interpretation

Evidence-backed count claims

Each row keeps its original unit and basis. Rows with different units or overlapping mappings are never added.

BenchmarkMapped taskCoverageCountVersionEvidence
BioMysteryBench
root: biomysterybench
End-to-end computational analysis
official-taxonomy · high
explicitly-in-scope90 problems
v11 mystery-bioinformatics problems after the June 2026 answer-key audit
v11biomysterybench-evidence-v11
BixBench
root: bixbench
End-to-end computational analysis
official-taxonomy · high
explicitly-in-scope205 questions
one question per row in the official v1.5 BixBench.jsonl
v1.5bixbench-evidence-v1-5-counts
BLADE
root: blade
End-to-end computational analysis
official-track · high
explicitly-in-scope12 problems
paired real-world research questions and datasets used as the source units for BLADE
arXiv v3blade-evidence-current-counts
BLADE End-to-End Analysis Generation
root: blade
End-to-end computational analysis
official-track · high
explicitly-in-scope12 problems
paired research questions and datasets requiring a complete generated analysis
arXiv v3blade-generation-evidence-counts
CompBioBench
root: compbiobench
End-to-end computational analysis
official-taxonomy · high
explicitly-in-scope100 tasks
v1 independent computational-biology tasks
v1compbiobench-evidence-counts
compbiobench-evidence-runner-license
GeneBench-Pro
root: genebench-pro
End-to-end computational analysis
official-taxonomy · high
explicitly-in-scope129 problems
self-contained synthetic scientific-analysis problems (called evaluations in the paper abstract)
paper-v1genebench-pro-paper-evidence
LAB-Bench
root: lab-bench
Experiment and protocol planning
official-track · high
explicitly-in-scope135 questions
ProtocolQA questions across public and private splits.
repository-998a8e0lab-bench-evidence-paper
LAB-Bench CloningScenarios
root: lab-bench
Experiment and protocol planning
official-track · high
explicitly-in-scope41 questions
questions across public and private splits
repository-998a8e0lab-bench-cloning-scenarios-evidence-paper
LAB-Bench ProtocolQA
root: lab-bench
Experiment and protocol planning
official-track · high
explicitly-in-scope135 questions
questions across public and private splits
repository-998a8e0lab-bench-protocolqa-evidence-paper
SCIGYM
root: scigym
Reaction-network reconstruction
official-taxonomy · high
explicitly-in-scope350 systems
distinct curated BioModels systems released as SBML benchmark instances
2025 releasescigym-evidence-release-counts
scigym-evidence-taxonomy
SCIGYM Large
root: scigym
Reaction-network reconstruction
official-track · high
explicitly-in-scope213 systems
unique SBML systems in the official large Parquet split, containing the remaining systems with up to 400 reactions
2025 releasescigym-large-evidence-count
SCIGYM Small
root: scigym
Reaction-network reconstruction
official-track · high
explicitly-in-scope137 systems
unique SBML systems with fewer than 10 reactions in the official small Parquet split
2025 releasescigym-small-evidence-count
LAB-Bench
root: lab-bench
Scientific database retrieval
official-track · high
explicitly-in-scope650 questions
Questions across the ten formal DbQA child tasks.
repository-998a8e0lab-bench-evidence-repository
LAB-Bench DbQA
root: lab-bench
Scientific database retrieval
official-track · high
explicitly-in-scope650 questions
questions across formal child tasks
repository-998a8e0lab-bench-dbqa-evidence-repository
LAB-Bench DbQA — Disease gene associations
root: lab-bench
Scientific database retrieval
official-track · high
explicitly-in-scope50 questions
questions across public and private splits
repository-998a8e0lab-bench-dbqa-dga-evidence-paper
LAB-Bench DbQA — Gene location
root: lab-bench
Scientific database retrieval
official-track · high
explicitly-in-scope50 questions
questions across public and private splits
repository-998a8e0lab-bench-dbqa-gene-location-evidence-paper
LAB-Bench DbQA — miRNA targets
root: lab-bench
Scientific database retrieval
official-track · high
explicitly-in-scope50 questions
questions across public and private splits
repository-998a8e0lab-bench-dbqa-mirna-targets-evidence-paper
LAB-Bench DbQA — Mouse tumor gene sets
root: lab-bench
Scientific database retrieval
official-track · high
explicitly-in-scope100 questions
questions across public and private splits
repository-998a8e0lab-bench-dbqa-mouse-tumor-gene-sets-evidence-paper
LAB-Bench DbQA — Oncogenic signatures
root: lab-bench
Scientific database retrieval
official-track · high
explicitly-in-scope50 questions
questions across public and private splits
repository-998a8e0lab-bench-dbqa-oncogenic-signatures-evidence-paper
LAB-Bench DbQA — GTRD transcription-factor binding sites
root: lab-bench
Scientific database retrieval
official-track · high
explicitly-in-scope50 questions
questions across public and private splits
repository-998a8e0lab-bench-dbqa-tfbs-gtrd-evidence-paper
LAB-Bench DbQA — Protein variant from sequence
root: lab-bench
Scientific database retrieval
official-track · high
explicitly-in-scope100 questions
questions across public and private splits
repository-998a8e0lab-bench-dbqa-variant-from-sequence-evidence-paper
LAB-Bench DbQA — Protein variant with multiple sequences
root: lab-bench
Scientific database retrieval
official-track · high
explicitly-in-scope100 questions
questions across public and private splits
repository-998a8e0lab-bench-dbqa-variant-multi-sequence-evidence-paper
LAB-Bench DbQA — Vaccine response gene sets
root: lab-bench
Scientific database retrieval
official-track · high
explicitly-in-scope50 questions
questions across public and private splits
repository-998a8e0lab-bench-dbqa-vax-response-evidence-paper
LAB-Bench DbQA — Viral protein–protein interactions
root: lab-bench
Scientific database retrieval
official-track · high
explicitly-in-scope50 questions
questions across public and private splits
repository-998a8e0lab-bench-dbqa-viral-ppi-evidence-paper
Anthropic Scientific Figure Interpretation Eval
root: anthropic-key-life-sciences-evals
Scientific evidence interpretation
official-track · high
explicitly-in-scopeNot reported
private scientific figure interpretation tasks
reported-2026-01-11
as of 2026-01-11
anthropic-scientific-figure-evidence
BLADE Decision-Discrimination MCQ
root: blade
Scientific evidence interpretation
official-track · high
explicitly-in-scope188 questions
individual multiple-choice decision-discrimination questions
arXiv v3blade-mcq-evidence-counts
LAB-Bench
root: lab-bench
Scientific evidence interpretation
official-track · high
explicitly-in-scopeNot reported
FigQA, LitQA2, SuppQA, and TableQA questions.
repository-998a8e0lab-bench-evidence-paper
LAB-Bench FigQA
root: lab-bench
Scientific evidence interpretation
official-track · high
explicitly-in-scope226 questions
questions across public and private splits
repository-998a8e0lab-bench-figqa-evidence-paper
LAB-Bench LitQA2
root: lab-bench
Scientific evidence interpretation
official-track · high
explicitly-in-scope248 questions
questions across public and private splits
repository-998a8e0lab-bench-litqa2-evidence-paper
LAB-Bench SuppQA
root: lab-bench
Scientific evidence interpretation
official-track · high
explicitly-in-scope102 questions
questions across public and private splits
repository-998a8e0lab-bench-suppqa-evidence-paper
LAB-Bench TableQA
root: lab-bench
Scientific evidence interpretation
official-track · high
explicitly-in-scope305 questions
questions across public and private splits
repository-998a8e0lab-bench-tableqa-evidence-paper
SCIGYM
root: scigym
Simulation-based experiment
official-taxonomy · high
explicitly-in-scope350 systems
distinct curated BioModels systems released as SBML benchmark instances
2025 releasescigym-evidence-release-counts
scigym-evidence-taxonomy
SCIGYM Large
root: scigym
Simulation-based experiment
official-track · high
explicitly-in-scope213 systems
unique SBML systems in the official large Parquet split, containing the remaining systems with up to 400 reactions
2025 releasescigym-large-evidence-count
SCIGYM Small
root: scigym
Simulation-based experiment
official-track · high
explicitly-in-scope137 systems
unique SBML systems with fewer than 10 reactions in the official small Parquet split
2025 releasescigym-small-evidence-count

Official evaluations connected to these benchmarks

Runs are included only for benchmark records mapped here (and formal child tracks when a mapped suite is the root). A task mapping does not imply that every run isolates this task.

WorkProvider / classRelated runs
Evaluating Claude's bioinformatics research capabilities with BioMysteryBenchAnthropic
benchmark_creator
biomysterybench-official-run
biomysterybench-v8-human-difficult
biomysterybench-v8-human-solvable
BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational BiologyFutureHouse, ScienceMachine
benchmark_creator
bixbench-creator-paper
bixbench-paper-mcq-no-images
bixbench-paper-mcq-no-refusal
bixbench-paper-mcq-refusal
BixBench v1.5 dataset and evaluation releaseFutureHouse, ScienceMachine
benchmark_creator
bixbench-v1-5-agentic-mcq-no-refusal-images
bixbench-v1-5-agentic-mcq-refusal-images
bixbench-v1-5-agentic-mcq-refusal-no-images
bixbench-v1-5-agentic-open-images
bixbench-v1-5-zero-shot-mcq-no-refusal
bixbench-v1-5-zero-shot-mcq-refusal
bixbench-v1-5-zero-shot-open
BLADE: Benchmarking Language Model Agents for Data-Driven ScienceUniversity of Washington, UC Berkeley, New York University, Stanford University, University of British Columbia, Microsoft, George Washington University
benchmark_creator
blade-creator-decision-mcq
blade-creator-paper
blade-creator-react
Agentic systems are adept at solving well-scoped, verifiable problems in computational biologyGenentech, Roche
benchmark_creator
compbiobench-codex-hardest
compbiobench-creator-full
compbiobench-haiku-full
compbiobench-haiku-hardest
compbiobench-nonagentic-baselines
compbiobench-opus-full
compbiobench-opus-hardest
compbiobench-sonnet-full
compbiobench-sonnet-hardest
GeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational BiomedicineOpenAI
benchmark_creator
genebench-pro-claude-high
genebench-pro-claude-low
genebench-pro-claude-max
genebench-pro-claude-medium
genebench-pro-claude-xhigh
genebench-pro-official
genebench-pro-pro-mode
genebench-pro-reasoning-enabled
genebench-pro-standard-high
genebench-pro-standard-low
genebench-pro-standard-max
genebench-pro-standard-medium
genebench-pro-standard-none
Claude Sonnet 4.5 System CardAnthropic
official_model_provider
lab-bench-cloning-scenarios-anthropic-sonnet45-system-card
lab-bench-figqa-anthropic-sonnet45-system-card
lab-bench-protocolqa-anthropic-sonnet45-system-card
lab-bench-seqqa-anthropic-sonnet45-system-card
LAB-Bench: Measuring Capabilities of Language Models for Biology ResearchFutureHouse
benchmark_creator
lab-bench-cloning-scenarios-creator-mcq
lab-bench-cloning-scenarios-creator-mcq-llama-context
lab-bench-cloning-scenarios-creator-open-response
lab-bench-dbqa-dga-creator-mcq
lab-bench-dbqa-gene-location-creator-mcq
lab-bench-dbqa-mirna-targets-creator-mcq
lab-bench-dbqa-mouse-tumor-gene-sets-creator-mcq
lab-bench-dbqa-oncogenic-signatures-creator-mcq
lab-bench-dbqa-tfbs-gtrd-creator-mcq
lab-bench-dbqa-variant-from-sequence-creator-mcq
lab-bench-dbqa-variant-multi-sequence-creator-mcq
lab-bench-dbqa-vax-response-creator-mcq
lab-bench-dbqa-viral-ppi-creator-mcq
lab-bench-figqa-creator-mcq
lab-bench-figqa-creator-open-response
lab-bench-litqa2-creator-mcq
lab-bench-protocolqa-creator-mcq
lab-bench-protocolqa-creator-open-response
lab-bench-suppqa-creator-mcq
lab-bench-tableqa-creator-mcq
Claude Sonnet 4.6 System CardAnthropic
official_model_provider
lab-bench-figqa-crop-tool
lab-bench-figqa-no-tools
Claude for Life SciencesAnthropic
official_model_provider
lab-bench-protocolqa-anthropic
Measuring Scientific Capabilities of Language Models with a Systems Biology Dry LabUniversity of Toronto, SickKids, Axiom, Mila, Vector Institute
benchmark_creator
scigym-small-creator-paper
scigym-small-zero-shot