Scientific domain

Life science

Broad applied and foundational life-science research.

生命科学

22 registered benchmarks and tracks

suite0 runs

AbBiBench

A framework using antibody–antigen complexes to evaluate affinity prediction and antibody redesign.

PredictionDesignGenerationOptimization
track1 runs

Anthropic Computational Biology Eval

Private Anthropic computational-biology evaluation direction reported only through a model-trend chart, without public tasks, counts, or protocol details.

Data analysisScientific reasoning
suite0 runs

Anthropic Key Life Sciences Evals

An Anthropic private internal suite reported only through an official accuracy chart covering scientific figure interpretation, computational biology, and protein understanding.

KnowledgeEvidence synthesisData analysisScientific reasoning
track1 runs

Anthropic Protein Understanding Eval

Private Anthropic protein-understanding evaluation direction reported only through a model-trend chart, without public tasks, counts, or protocol details.

KnowledgeScientific reasoning
track1 runs

Anthropic Scientific Figure Interpretation Eval

Private Anthropic evaluation direction for scientific figure interpretation, reported only through a model-trend chart with no task count or released examples.

Evidence synthesisScientific reasoning
suite0 runs

BLADE

A cross-domain suite for discerning defensible analysis decisions and generating executable end-to-end analyses for open-ended scientific research questions; four of its twelve source questions are explicitly biological or ecological.

KnowledgeClassificationData analysisCoding
track2 runs

BLADE End-to-End Analysis Generation

The BLADE track requiring a conceptual-variable specification, executable data-transformation function, and statistical-model function for each open-ended research question and dataset.

Data analysisCodingTool useScientific reasoning
track1 runs

BLADE Decision-Discrimination MCQ

The BLADE track for selecting the most or least justifiable conceptual-variable and data-transformation decisions for a research question and dataset.

KnowledgeClassificationData analysisScientific reasoning
dataset0 runs

CaM benchmark

A multistate protein sequence-design benchmark spanning CaM conformations and binding modes.

DesignOptimization
suite0 runs

crafted experiments

Real single-cell RNA-seq data augmented with known gene perturbations for comparing feature-selection methods.

Data analysis
suite0 runs

LAB-Bench

A practical biology-research suite of 2,457 multiple-choice questions across eight broad categories and 31 versioned task files, with public and private contamination-monitoring splits.

KnowledgeEvidence synthesisRetrievalPrediction
track5 runs

LAB-Bench FigQA

Multiple-choice interpretation and multi-element reasoning over scientific figures shown without captions or paper context.

Evidence synthesisScientific reasoning
track1 runs

LAB-Bench LitQA2

Literature-retrieval questions whose answers require findings in full research papers rather than titles or abstracts.

RetrievalEvidence synthesisScientific reasoning
track1 runs

LAB-Bench SuppQA

Retrieval and interpretation questions answerable from paper supplementary text or PDF tables.

RetrievalEvidence synthesisScientific reasoning
track1 runs

LAB-Bench TableQA

Lookup, calculation, and reasoning questions over table images extracted from scientific papers.

Evidence synthesisData analysisScientific reasoning
agentic-eval1 runs

LifeSciBench

Expert-authored, artifact-rich free-response tasks that evaluate realistic research judgment across applied life-science workflows.

Evidence synthesisRetrievalDesignGeneration
dataset0 runs

PapD benchmark

A multistate protein sequence-design benchmark targeting the multispecific PapD binding interface.

DesignOptimization
dataset0 runs

RfaH benchmark

A multistate protein sequence-design benchmark using the fold-switching conformations of RfaH.

DesignOptimization
agentic-eval0 runs

scBench

Agentic evaluation suite for data-grounded single-cell analysis across diverse sequencing technologies and workflow stages.

Data analysisCodingTool useScientific reasoning
suite0 runs

SCIGYM

An agentic systems-biology suite in which language models iteratively perturb simulated SBML systems, analyze time-series observations in Python, and reconstruct hidden biological reactions.

Experiment planningData analysisCodingTool use
track0 runs

SCIGYM Large

The formally released SCIGYM track containing the 213 systems not included in the creator paper's model evaluation, with systems reaching up to 400 reactions.

Experiment planningData analysisCodingTool use
track2 runs

SCIGYM Small

The formally released and creator-evaluated SCIGYM track containing biological systems with fewer than ten reactions.

Experiment planningData analysisCodingTool use

Capability coverage

Registry records tagged Life science, counted by capability.

CSV ↓
Accessible data table
CapabilityRecords
Knowledge5
Evidence synthesis8
Retrieval4
Prediction2
Classification3
Design6
Generation2
Optimization5
Data analysis13
Coding6
Tool use7
Experiment planning5
Troubleshooting2
Scientific reasoning17
Scientific communication1
Coverage-gap reading: zero counts indicate a gap in this registry, not proof that no benchmark exists. Propose a primary source through the contribution form.