Scientific domain

Protein structure

Protein folding

蛋白质结构

15 registered benchmarks and tracks

suite1 runs

ATOM3D

A living collection of eight curated 3D molecular-learning tasks spanning small molecules, protein interactions and mutations, ligand binding, and protein/RNA structure ranking.

PredictionClassificationRegression
agentic-eval3 runs

BioMysteryBench

An agentic bioinformatics benchmark of objective, expert-authored mysteries over anonymized real-world biological data, scored on final answers rather than prescribed analysis paths.

Data analysisCodingTool useScientific reasoning
dataset0 runs

CaM benchmark

A multistate protein sequence-design benchmark spanning CaM conformations and binding modes.

DesignOptimization
competition3 runs

CAMEO

Weekly, automated, independent, blind evaluation of registered macromolecular structure-prediction servers on complete PDB entries whose experimental structures are withheld during prediction.

Prediction
competition0 runs

CASP

Biennial blind community experiments that assess macromolecular structure, complex, ligand, and model-accuracy prediction against experimental structures withheld during prediction.

Prediction
track0 runs

CASP17 Immune Complexes

Dedicated CASP17 category for blind prediction of antibody-antigen, nanobody-antigen, and T-cell receptor complex structures.

Prediction
track3 runs

CASP Protein-Ligand Prediction

Formal CASP track for blind prediction of protein-ligand binding poses, binding affinity or rank, binding pockets, and pose confidence.

PredictionRegression
track1 runs

CASP Protein Monomers

Formal CASP track assessing blind predictions of single-protein structures and post hoc protein evaluation units against withheld experimental coordinates.

Prediction
track1 runs

CASP Protein Multimers

Formal CASP track assessing blind protein-complex predictions, including overall folds, interfaces, stoichiometry-free phases, and model-selection phases.

Prediction
agentic-eval9 runs

CompBioBench

A 100-task agent benchmark of objectively gradable computational-biology problems requiring multi-step reasoning, bespoke code, tools, and real-world external resources.

Data analysisCodingTool useRetrieval
agentic-eval1 runs

LifeSciBench

Expert-authored, artifact-rich free-response tasks that evaluate realistic research judgment across applied life-science workflows.

Evidence synthesisRetrievalDesignGeneration
dataset0 runs

PapD benchmark

A multistate protein sequence-design benchmark targeting the multispecific PapD binding interface.

DesignOptimization
dataset1 runs

ProteinLMBench

A creator-curated set of 944 protein-science multiple-choice questions with answer explanations, generated from research literature and released for evaluating text LLM protein understanding.

KnowledgeScientific reasoningPrediction
dataset0 runs

RfaH benchmark

A multistate protein sequence-design benchmark using the fold-switching conformations of RfaH.

DesignOptimization
suite1 runs

TAPE

A five-task benchmark for protein representation learning spanning secondary structure, residue contacts, remote homology, fluorescence, and stability.

PredictionClassificationRegression

Capability coverage

Registry records tagged Protein structure, counted by capability.

CSV ↓
Accessible data table
CapabilityRecords
Knowledge1
Evidence synthesis1
Retrieval2
Prediction9
Classification2
Regression3
Design4
Generation1
Optimization4
Data analysis3
Coding2
Tool use2
Experiment planning1
Troubleshooting1
Scientific reasoning4
Scientific communication1
Coverage-gap reading: zero counts indicate a gap in this registry, not proof that no benchmark exists. Propose a primary source through the contribution form.