Scientific domain

Protein-protein binding

Protein-protein interaction

蛋白-蛋白结合

19 registered benchmarks and tracks

dataset0 runs

AB-Bind

Antibody-focused mutational binding data with accompanying structures for computational affinity prediction.

PredictionClassificationRegressionOptimization
suite0 runs

AbBiBench

A framework using antibody–antigen complexes to evaluate affinity prediction and antibody redesign.

PredictionDesignGenerationOptimization
suite1 runs

ATOM3D

A living collection of eight curated 3D molecular-learning tasks spanning small molecules, protein interactions and mutations, ligand binding, and protein/RNA structure ranking.

PredictionClassificationRegression
suite0 runs

Biology-Instructions

A multi-omics sequence-instruction suite with 21 formal predictive tasks across DNA, RNA, protein, and multi-molecule inputs; the final paper reports task-specific held-out evaluations rather than a single aggregate score.

PredictionClassificationRegressionScientific reasoning
dataset0 runs

CaM benchmark

A multistate protein sequence-design benchmark spanning CaM conformations and binding modes.

DesignOptimization
competition3 runs

CAMEO

Weekly, automated, independent, blind evaluation of registered macromolecular structure-prediction servers on complete PDB entries whose experimental structures are withheld during prediction.

Prediction
competition0 runs

CASP

Biennial blind community experiments that assess macromolecular structure, complex, ligand, and model-accuracy prediction against experimental structures withheld during prediction.

Prediction
track0 runs

CASP17 Immune Complexes

Dedicated CASP17 category for blind prediction of antibody-antigen, nanobody-antigen, and T-cell receptor complex structures.

Prediction
track1 runs

CASP Protein Multimers

Formal CASP track assessing blind protein-complex predictions, including overall folds, interfaces, stoichiometry-free phases, and model-selection phases.

Prediction
suite0 runs

FLIP

A supervised protein sequence-to-fitness benchmark that turns three experimental landscapes into 15 biologically motivated dataset splits for testing generalization in protein engineering.

PredictionRegression
track5 runs

FLIP GB1

Five supervised splits over a downsampled, highly epistatic four-site GB1 immunoglobulin-binding landscape, designed to test mutation-depth and low-to-high-fitness generalization.

PredictionRegression
suite0 runs

LAB-Bench

A practical biology-research suite of 2,457 multiple-choice questions across eight broad categories and 31 versioned task files, with public and private contamination-monitoring splits.

KnowledgeEvidence synthesisRetrievalPrediction
track0 runs

LAB-Bench DbQA

Database-retrieval category spanning 10 genomics, clinical, protein, regulatory, vaccine-response, and viral-PPI tasks.

RetrievalKnowledgePredictionClassification
agentic-eval1 runs

LifeSciBench

Expert-authored, artifact-rich free-response tasks that evaluate realistic research judgment across applied life-science workflows.

Evidence synthesisRetrievalDesignGeneration
dataset0 runs

PapD benchmark

A multistate protein sequence-design benchmark targeting the multispecific PapD binding interface.

DesignOptimization
dataset0 runs

PPB-Affinity

A reusable protein-protein binding-affinity dataset with complex structures, measured affinities, receptor and ligand chains, and mutation annotations.

Prediction
dataset1 runs

ProteinLMBench

A creator-curated set of 944 protein-science multiple-choice questions with answer explanations, generated from research literature and released for evaluating text LLM protein understanding.

KnowledgeScientific reasoningPrediction

Capability coverage

Registry records tagged Protein-protein binding, counted by capability.

CSV ↓
Accessible data table
CapabilityRecords
Knowledge4
Evidence synthesis2
Retrieval4
Prediction16
Classification6
Regression5
Design5
Generation2
Optimization5
Data analysis2
Tool use1
Experiment planning2
Troubleshooting2
Scientific reasoning5
Scientific communication1
Coverage-gap reading: zero counts indicate a gap in this registry, not proof that no benchmark exists. Propose a primary source through the contribution form.