AbBiBench
A framework using antibody–antigen complexes to evaluate affinity prediction and antibody redesign.
Scientific domain
Broad applied and foundational life-science research.
生命科学
A framework using antibody–antigen complexes to evaluate affinity prediction and antibody redesign.
Private Anthropic computational-biology evaluation direction reported only through a model-trend chart, without public tasks, counts, or protocol details.
An Anthropic private internal suite reported only through an official accuracy chart covering scientific figure interpretation, computational biology, and protein understanding.
Private Anthropic protein-understanding evaluation direction reported only through a model-trend chart, without public tasks, counts, or protocol details.
Private Anthropic evaluation direction for scientific figure interpretation, reported only through a model-trend chart with no task count or released examples.
A cross-domain suite for discerning defensible analysis decisions and generating executable end-to-end analyses for open-ended scientific research questions; four of its twelve source questions are explicitly biological or ecological.
The BLADE track requiring a conceptual-variable specification, executable data-transformation function, and statistical-model function for each open-ended research question and dataset.
The BLADE track for selecting the most or least justifiable conceptual-variable and data-transformation decisions for a research question and dataset.
A multistate protein sequence-design benchmark spanning CaM conformations and binding modes.
Real single-cell RNA-seq data augmented with known gene perturbations for comparing feature-selection methods.
A practical biology-research suite of 2,457 multiple-choice questions across eight broad categories and 31 versioned task files, with public and private contamination-monitoring splits.
Multiple-choice interpretation and multi-element reasoning over scientific figures shown without captions or paper context.
Literature-retrieval questions whose answers require findings in full research papers rather than titles or abstracts.
Retrieval and interpretation questions answerable from paper supplementary text or PDF tables.
Lookup, calculation, and reasoning questions over table images extracted from scientific papers.
Expert-authored, artifact-rich free-response tasks that evaluate realistic research judgment across applied life-science workflows.
A multistate protein sequence-design benchmark targeting the multispecific PapD binding interface.
A multistate protein sequence-design benchmark using the fold-switching conformations of RfaH.
Agentic evaluation suite for data-grounded single-cell analysis across diverse sequencing technologies and workflow stages.
An agentic systems-biology suite in which language models iteratively perturb simulated SBML systems, analyze time-series observations in Python, and reconstruct hidden biological reactions.
The formally released SCIGYM track containing the 213 systems not included in the creator paper's model evaluation, with systems reaching up to 400 reactions.
The formally released and creator-evaluated SCIGYM track containing biological systems with fewer than ten reactions.
Registry records tagged Life science, counted by capability.
| Capability | Records |
|---|---|
| Knowledge | 5 |
| Evidence synthesis | 8 |
| Retrieval | 4 |
| Prediction | 2 |
| Classification | 3 |
| Design | 6 |
| Generation | 2 |
| Optimization | 5 |
| Data analysis | 13 |
| Coding | 6 |
| Tool use | 7 |
| Experiment planning | 5 |
| Troubleshooting | 2 |
| Scientific reasoning | 17 |
| Scientific communication | 1 |
Task mappings are evidence-backed and may be partial for mixed suites.