Scientific domain

Protein sequence

Amino-acid sequence understanding and modeling.

蛋白质序列

32 registered benchmarks and tracks

suite0 runs

Biology-Instructions

A multi-omics sequence-instruction suite with 21 formal predictive tasks across DNA, RNA, protein, and multi-molecule inputs; the final paper reports task-specific held-out evaluations rather than a single aggregate score.

PredictionClassificationRegressionScientific reasoning
dataset0 runs

CaM benchmark

A multistate protein sequence-design benchmark spanning CaM conformations and binding modes.

DesignOptimization
suite0 runs

FLIP

A supervised protein sequence-to-fitness benchmark that turns three experimental landscapes into 15 biologically motivated dataset splits for testing generalization in protein engineering.

PredictionRegression
track7 runs

FLIP AAV

Seven supervised splits over sampled and machine-designed AAV2 VP-1 capsid variants, measuring generalization across mutation depth, fitness, and sampled-versus-designed pools.

PredictionRegression
track5 runs

FLIP GB1

Five supervised splits over a downsampled, highly epistatic four-site GB1 immunoglobulin-binding landscape, designed to test mutation-depth and low-to-high-fitness generalization.

PredictionRegression
track3 runs

FLIP Meltome Thermostability

Three supervised sequence-to-melting-temperature splits spanning all species, human proteins, and a single human cell line, with sequence-cluster-aware train/test separation.

PredictionRegression
suite0 runs

LAB-Bench

A practical biology-research suite of 2,457 multiple-choice questions across eight broad categories and 31 versioned task files, with public and private contamination-monitoring splits.

KnowledgeEvidence synthesisRetrievalPrediction
track0 runs

LAB-Bench DbQA

Database-retrieval category spanning 10 genomics, clinical, protein, regulatory, vaccine-response, and viral-PPI tasks.

RetrievalKnowledgePredictionClassification
track1 runs

LAB-Bench SeqQA

Sequence-comprehension and manipulation category spanning 15 formal tasks involving PCR, restriction digestion, ORFs, translation, GC content, and DNA–protein relationships.

Data analysisDesignPredictionScientific reasoning
agentic-eval1 runs

LifeSciBench

Expert-authored, artifact-rich free-response tasks that evaluate realistic research judgment across applied life-science workflows.

Evidence synthesisRetrievalDesignGeneration
dataset0 runs

PapD benchmark

A multistate protein sequence-design benchmark targeting the multispecific PapD binding interface.

DesignOptimization
suite0 runs

ProteinGym

Versioned deep-mutational-scanning and clinical-variant benchmarks for protein fitness prediction and design in zero-shot and supervised regimes.

PredictionRegressionClassificationDesign
track0 runs

ProteinGym Clinical Indels

ProteinGym track for classifying short human clinical insertion and deletion variants against ClinVar and gnomAD-derived labels.

PredictionClassification
track0 runs

ProteinGym Clinical Substitutions

ProteinGym track for classifying expert-annotated human clinical substitution variants on a per-protein basis.

PredictionClassification
track0 runs

ProteinGym DMS Indels

ProteinGym track for predicting experimental fitness measurements of insertion and deletion mutants across deep-mutational-scanning assays.

PredictionRegressionClassificationDesign
track1 runs

ProteinGym DMS Substitutions

ProteinGym track for predicting experimental fitness measurements of substitution mutants across deep-mutational-scanning assays.

PredictionRegressionClassificationDesign
dataset1 runs

ProteinLMBench

A creator-curated set of 944 protein-science multiple-choice questions with answer explanations, generated from research literature and released for evaluating text LLM protein understanding.

KnowledgeScientific reasoningPrediction
dataset0 runs

RfaH benchmark

A multistate protein sequence-design benchmark using the fold-switching conformations of RfaH.

DesignOptimization
suite1 runs

TAPE

A five-task benchmark for protein representation learning spanning secondary structure, residue contacts, remote homology, fluorescence, and stability.

PredictionClassificationRegression

Capability coverage

Registry records tagged Protein sequence, counted by capability.

CSV ↓
Accessible data table
CapabilityRecords
Knowledge3
Evidence synthesis2
Retrieval5
Prediction23
Classification15
Regression12
Design9
Generation1
Optimization7
Data analysis6
Tool use2
Experiment planning2
Troubleshooting2
Scientific reasoning12
Scientific communication1
Coverage-gap reading: zero counts indicate a gap in this registry, not proof that no benchmark exists. Propose a primary source through the contribution form.