AB-Bind
Antibody-focused mutational binding data with accompanying structures for computational affinity prediction.
Registry explorer
Search the scientific scope, filter access and task type, then inspect every work and run behind an entry. Filters persist in the URL.
109 records · Download all benchmarks CSV
| Benchmark | Kind / release | Domain | Scientific task | Access | Runs |
|---|---|---|---|---|---|
| AB-Bind Antibody-focused mutational binding data with accompanying structures for computational affinity prediction. | dataset 2015-11-06 | Fully open | 0 | ||
| AbBiBench A framework using antibody–antigen complexes to evaluate affinity prediction and antibody redesign. | suite 2025-05-23 | Fully open Provisional · medium | 0 | ||
| Anthropic Computational Biology Eval Private Anthropic computational-biology evaluation direction reported only through a model-trend chart, without public tasks, counts, or protocol details. | track 2026-01-11 | Private or internal | 1 | ||
| Anthropic Key Life Sciences Evals An Anthropic private internal suite reported only through an official accuracy chart covering scientific figure interpretation, computational biology, and protein understanding. | suite 2026-01-11 | Private or internal | 0 | ||
| Anthropic Protein Understanding Eval Private Anthropic protein-understanding evaluation direction reported only through a model-trend chart, without public tasks, counts, or protocol details. | track 2026-01-11 | Private or internal | 1 | ||
| Anthropic Scientific Figure Interpretation Eval Private Anthropic evaluation direction for scientific figure interpretation, reported only through a model-trend chart with no task count or released examples. | track 2026-01-11 | Private or internal | 1 | ||
| ATOM3D A living collection of eight curated 3D molecular-learning tasks spanning small molecules, protein interactions and mutations, ligand binding, and protein/RNA structure ranking. | suite 2020-12-07 | Fully open | 1 | ||
| BEACON A 13-task RNA representation benchmark covering structure, function, processing, modification, translation, degradation, programmable switches, and CRISPR activity. | suite 2024-06-14 | Fully open | 1 | ||
| benchmark-dual human CRISPR-Cas9 library A human CRISPR-Cas9 paired-guide library created to compare dual- and single-targeting strategies in loss-of-function screens. | dataset 2025-02-26 | Fully open | 0 | ||
| benchmark human CRISPR-Cas9 library A human CRISPR-Cas9 guide-RNA library assembled to compare single-targeting library performance in loss-of-function screens. | dataset 2025-02-26 | Fully open | 0 | ||
| Biology-Instructions A multi-omics sequence-instruction suite with 21 formal predictive tasks across DNA, RNA, protein, and multi-molecule inputs; the final paper reports task-specific held-out evaluations rather than a single aggregate score. | suite 2024-12-26 | Partially open | 0 | ||
| Biology-Instructions Antibody-Antigen Neutralization Binary neutralization prediction for an antibody-antigen protein-sequence pair, evaluated with Matthews correlation coefficient. | track 2024-12-26 | Partially open | 3 | ||
| Biology-Instructions APA Isoform Prediction Regression of alternative-polyadenylation isoform usage from an RNA sequence, reported as squared Pearson correlation under the paper's R2 label. | track 2024-12-26 | Partially open | 3 | ||
| Biology-Instructions Core Promoter Detection Binary detection of a core promoter in a short DNA sequence, evaluated with Matthews correlation coefficient. | track 2024-12-26 | Partially open | 3 | ||
| Biology-Instructions CRISPR On-Target Prediction Regression of CRISPR guide on-target activity from an RNA sequence, evaluated with Spearman rank correlation. | track 2024-12-26 | Partially open | 3 | ||
| Biology-Instructions Enhancer Activity Prediction Two-output regression of housekeeping and developmental enhancer activity from a DNA sequence, evaluated with separate Pearson correlations. | track 2024-12-26 | Partially open | 3 | ||
| Biology-Instructions Enzyme Commission Number Prediction Multi-label Enzyme Commission number prediction from a protein sequence, evaluated with the creator's Fmax implementation. | track 2024-12-26 | Partially open | 3 | ||
| Biology-Instructions Epigenetic Marks Prediction Binary prediction of whether a DNA sequence carries an epigenetic mark, evaluated with Matthews correlation coefficient. | track 2024-12-26 | Partially open | 3 | ||
| Biology-Instructions Enhancer-Promoter Interaction Prediction Binary interaction prediction for enhancer and promoter DNA sequences, evaluated with Matthews correlation coefficient. | track 2024-12-26 | Partially open | 3 | ||
| Biology-Instructions Protein Fluorescence Prediction Regression of protein fluorescence from an amino-acid sequence, evaluated with Spearman rank correlation. | track 2024-12-26 | Partially open | 3 | ||
| Biology-Instructions RNA Modification Prediction Multi-label prediction of RNA chemical modifications, evaluated with macro area under the ROC curve. | track 2024-12-26 | Partially open | 3 | ||
| Biology-Instructions Mean Ribosome Loading Prediction Regression of mean ribosome loading from an RNA sequence, reported as squared Pearson correlation under the paper's R2 label. | track 2024-12-26 | Partially open | 3 | ||
| Biology-Instructions Non-coding RNA Function Classification Thirteen-class functional classification of a non-coding RNA sequence, evaluated with exact extracted-label accuracy. | track 2024-12-26 | Partially open | 3 | ||
| Biology-Instructions Promoter Detection 300 Binary promoter detection in a 300-base-pair DNA context, evaluated with Matthews correlation coefficient. | track 2024-12-26 | Partially open | 3 | ||
| Biology-Instructions Programmable RNA Switches Three-output regression of ON, OFF, and ON/OFF programmable RNA-switch values, aggregated as their mean squared Pearson correlation. | track 2024-12-26 | Partially open | 3 | ||
| Biology-Instructions RNA-Protein Interaction Prediction Binary interaction prediction for an RNA and protein sequence pair, evaluated with Matthews correlation coefficient. | track 2024-12-26 | Partially open | 3 | ||
| Biology-Instructions siRNA Efficiency Prediction Regression of siRNA efficiency from paired sequence context, evaluated with the SAIS-inspired mixed score. | track 2024-12-26 | Partially open | 3 | ||
| Biology-Instructions Protein Solubility Prediction Binary protein-solubility prediction from an amino-acid sequence, evaluated with accuracy. | track 2024-12-26 | Partially open | 3 | ||
| Biology-Instructions Protein Stability Prediction Regression of protein stability from an amino-acid sequence, evaluated with Spearman rank correlation. | track 2024-12-26 | Partially open | 3 | ||
| Biology-Instructions Human Transcription Binding Sites Detection Binary detection of transcription-factor binding sites in human DNA sequences, evaluated with Matthews correlation coefficient. | track 2024-12-26 | Partially open | 3 | ||
| Biology-Instructions Mouse Transcription Binding Sites Detection Binary detection of transcription-factor binding sites in mouse DNA sequences, evaluated with Matthews correlation coefficient. | track 2024-12-26 | Partially open | 3 | ||
| Biology-Instructions Protein Thermostability Prediction Regression of protein thermostability from an amino-acid sequence, evaluated with Spearman rank correlation. | track 2024-12-26 | Partially open | 3 | ||
| BioMysteryBench An agentic bioinformatics benchmark of objective, expert-authored mysteries over anonymized real-world biological data, scored on final answers rather than prescribed analysis paths. | agentic-eval 2026-04-29 | Partially open | 3 | ||
| BioSecBench-Surveillance Agentic evaluation of pathogen genomic-surveillance workflow selection and analysis from raw or near-raw sequencing data. | agentic-eval 2026-07-21 | Partially open | 3 | ||
| BixBench A containerized benchmark of long-horizon bioinformatics analysis over real published notebooks and associated data, with open-answer and multiple-choice evaluation modes. | agentic-eval 2025-02-28 | Fully open | 11 | ||
| BLADE A cross-domain suite for discerning defensible analysis decisions and generating executable end-to-end analyses for open-ended scientific research questions; four of its twelve source questions are explicitly biological or ecological. | suite 2024-08-19 | Fully open | 0 | ||
| BLADE End-to-End Analysis Generation The BLADE track requiring a conceptual-variable specification, executable data-transformation function, and statistical-model function for each open-ended research question and dataset. | track 2024-08-19 | Fully open | 2 | ||
| BLADE Decision-Discrimination MCQ The BLADE track for selecting the most or least justifiable conceptual-variable and data-transformation decisions for a research question and dataset. | track 2024-08-19 | Fully open | 1 | ||
| CaM benchmark A multistate protein sequence-design benchmark spanning CaM conformations and binding modes. | dataset 2024-07-11 | Fully open | 0 | ||
| CAMEO Weekly, automated, independent, blind evaluation of registered macromolecular structure-prediction servers on complete PDB entries whose experimental structures are withheld during prediction. | competition 2012-01-01 | Partially open | 3 | ||
| CASP Biennial blind community experiments that assess macromolecular structure, complex, ligand, and model-accuracy prediction against experimental structures withheld during prediction. | competition 1994-01-01 | Partially open | 0 | ||
| CASP17 Immune Complexes Dedicated CASP17 category for blind prediction of antibody-antigen, nanobody-antigen, and T-cell receptor complex structures. | track 2026-04-28 | Partially open | 0 | ||
| CASP Protein-Ligand Prediction Formal CASP track for blind prediction of protein-ligand binding poses, binding affinity or rank, binding pockets, and pose confidence. | track 2024-05-01 | Partially open | 3 | ||
| CASP Protein Monomers Formal CASP track assessing blind predictions of single-protein structures and post hoc protein evaluation units against withheld experimental coordinates. | track 1994-01-01 | Partially open | 1 | ||
| CASP Protein Multimers Formal CASP track assessing blind protein-complex predictions, including overall folds, interfaces, stoichiometry-free phases, and model-selection phases. | track 2024-05-01 | Partially open | 1 | ||
| CompBioBench A 100-task agent benchmark of objectively gradable computational-biology problems requiring multi-step reasoning, bespoke code, tools, and real-world external resources. | agentic-eval 2026-04-06 | Partially open | 9 | ||
| Comprehensive benchmark of differential transcript usage analysis for bulk and single-cell RNA sequencing A reusable benchmark for comparing differential transcript usage detection tools across simulated and real transcriptomics data. | suite Provisional · medium 2025-09-11 | Fully open Provisional · medium | 0 | ||
| crafted experiments Real single-cell RNA-seq data augmented with known gene perturbations for comparing feature-selection methods. | suite 2025-01-07 | Fully open | 0 | ||
| FLIP A supervised protein sequence-to-fitness benchmark that turns three experimental landscapes into 15 biologically motivated dataset splits for testing generalization in protein engineering. | suite 2021-10-11 | Fully open | 0 | ||
| FLIP AAV Seven supervised splits over sampled and machine-designed AAV2 VP-1 capsid variants, measuring generalization across mutation depth, fitness, and sampled-versus-designed pools. | track 2021-10-11 | Fully open | 7 | ||
| FLIP GB1 Five supervised splits over a downsampled, highly epistatic four-site GB1 immunoglobulin-binding landscape, designed to test mutation-depth and low-to-high-fitness generalization. | track 2021-10-11 | Fully open | 5 | ||
| FLIP Meltome Thermostability Three supervised sequence-to-melting-temperature splits spanning all species, human proteins, and a single human cell line, with sequence-cluster-aware train/test separation. | track 2021-10-11 | Fully open | 3 | ||
| GeneBench-Pro A research-level agent benchmark of 129 synthetic, multistage computational-biology analyses that require iterative QC, statistical modeling, diagnostics, and decision-relevant judgment. | agentic-eval 2026-06-30 | Partially open | 13 | ||
| Genomic Benchmarks A versioned collection of nine DNA sequence-classification datasets covering regulatory elements, promoters, enhancers, open chromatin, species, and coding-context discrimination. | suite 2023-05-01 | Fully open | 1 | ||
| GuacaMol A reproducible benchmark for de novo molecular design with five distribution-learning tests and twenty goal-directed generation problems in the current v2 suite. | suite 2018-11-23 | Fully open | 1 | ||
| LAB-Bench A practical biology-research suite of 2,457 multiple-choice questions across eight broad categories and 31 versioned task files, with public and private contamination-monitoring splits. | suite 2024-07-14 | Partially open | 0 | ||
| LAB-Bench CloningScenarios Human-hard, multi-step multiple-choice scenarios involving plasmids, DNA fragments, enzymes, and molecular-cloning workflows. | track 2024-07-14 | Partially open | 4 | ||
| LAB-Bench DbQA Database-retrieval category spanning 10 genomics, clinical, protein, regulatory, vaccine-response, and viral-PPI tasks. | track 2024-07-14 | Partially open | 0 | ||
| LAB-Bench DbQA — Disease gene associations Identifies genes associated with a phenotype in DisGeNET but not OMIM. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench DbQA — Gene location Retrieves human-gene cytogenetic locations from the stated Ensembl release. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench DbQA — miRNA targets Retrieves computationally predicted human miRNA targets from miRDB. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench DbQA — Mouse tumor gene sets Retrieves genes in Mammalian Phenotype Tumor Ontology gene sets. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench DbQA — Oncogenic signatures Retrieves membership in MSigDB C6 oncogenic-signature gene sets. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench DbQA — GTRD transcription-factor binding sites Retrieves promoter-region transcription-factor binding-site annotations from GTRD. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench DbQA — Protein variant from sequence Uses a protein sequence and ClinVar lookup to identify benign or pathogenic variants. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench DbQA — Protein variant with multiple sequences Identifies ClinVar variant pathogenicity while reasoning across multiple protein sequences. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench DbQA — Vaccine response gene sets Retrieves membership in MSigDB vaccine-response gene sets. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench DbQA — Viral protein–protein interactions Retrieves predicted human interaction partners of viral proteins from P-HIPSter. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench FigQA Multiple-choice interpretation and multi-element reasoning over scientific figures shown without captions or paper context. | track 2024-07-14 | Partially open | 5 | ||
| LAB-Bench LitQA2 Literature-retrieval questions whose answers require findings in full research papers rather than titles or abstracts. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench ProtocolQA Troubleshoots intentionally modified published biological protocols by selecting steps that would repair the stated outcome. | track 2024-07-14 | Partially open | 4 | ||
| LAB-Bench SeqQA Sequence-comprehension and manipulation category spanning 15 formal tasks involving PCR, restriction digestion, ORFs, translation, GC content, and DNA–protein relationships. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench SeqQA — ORF amino-acid position Finds the amino acid encoded at a specified position in the longest ORF of a DNA sequence. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench SeqQA — ORF amino-acid sequence Translates the longest ORF in a DNA sequence to its amino-acid sequence. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench SeqQA — ORF count above length Counts open reading frames encoding proteins above a specified amino-acid length. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench SeqQA — Translation efficiency Selects an RNA sequence whose ORF context is most likely to yield high translation efficiency. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench SeqQA — Gene-to-restriction primers Selects restriction-cloning primers from a named gene and enzyme pair. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench SeqQA — Gene-to-Gibson primers (HindIII) Selects primers for Gibson assembly into a HindIII-linearized vector. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench SeqQA — Gene-to-Gibson primers (SmaI) Selects primers for Gibson assembly into a SmaI-linearized vector. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench SeqQA — Primers-to-restriction enzymes Infers restriction enzymes from a gene name and primer pair. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench SeqQA — Amplicon length to primers Selects primers that produce a requested amplicon length from a DNA template. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench SeqQA — Primers to amplicon length Calculates expected amplicon length from a primer pair and DNA template. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench SeqQA — Sequence-to-restriction primers Selects restriction-cloning primers from an explicit gene sequence and enzyme pair. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench SeqQA — Amplicon sequence to primers Selects primers that produce a requested amplicon sequence from a DNA template. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench SeqQA — GC percentage Calculates the rounded GC percentage of a DNA sequence. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench SeqQA — Restriction-fragment lengths Calculates fragment lengths after restriction digestion of a DNA sequence. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench SeqQA — Restriction-fragment count Calculates the number of fragments after restriction digestion of a DNA sequence. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench SuppQA Retrieval and interpretation questions answerable from paper supplementary text or PDF tables. | track 2024-07-14 | Partially open | 1 | ||
| LAB-Bench TableQA Lookup, calculation, and reasoning questions over table images extracted from scientific papers. | track 2024-07-14 | Partially open | 1 | ||
| LifeSciBench Expert-authored, artifact-rich free-response tasks that evaluate realistic research judgment across applied life-science workflows. | agentic-eval 2026-06-17 | Private or internal | 1 | ||
| MoleculeNet The original molecular-machine-learning benchmark of 17 dataset collections and more than 800 prediction endpoints spanning quantum, physicochemical, biophysical, and physiological properties. | suite 2017-03-02 | Fully open | 1 | ||
| PapD benchmark A multistate protein sequence-design benchmark targeting the multispecific PapD binding interface. | dataset 2024-07-11 | Fully open | 0 | ||
| PPB-Affinity A reusable protein-protein binding-affinity dataset with complex structures, measured affinities, receptor and ligand chains, and mutation annotations. | dataset 2024-12-03 | Fully open | 0 | ||
| ProteinGym Versioned deep-mutational-scanning and clinical-variant benchmarks for protein fitness prediction and design in zero-shot and supervised regimes. | suite 2023-12-08 | Fully open | 0 | ||
| ProteinGym Clinical Indels ProteinGym track for classifying short human clinical insertion and deletion variants against ClinVar and gnomAD-derived labels. | track 2023-12-08 | Fully open | 0 | ||
| ProteinGym Clinical Substitutions ProteinGym track for classifying expert-annotated human clinical substitution variants on a per-protein basis. | track 2023-12-08 | Fully open | 0 | ||
| ProteinGym DMS Indels ProteinGym track for predicting experimental fitness measurements of insertion and deletion mutants across deep-mutational-scanning assays. | track 2023-12-08 | Fully open | 0 | ||
| ProteinGym DMS Substitutions ProteinGym track for predicting experimental fitness measurements of substitution mutants across deep-mutational-scanning assays. | track 2023-12-08 | Fully open | 1 | ||
| ProteinLMBench A creator-curated set of 944 protein-science multiple-choice questions with answer explanations, generated from research literature and released for evaluating text LLM protein understanding. | dataset 2024-04-29 | Fully open | 1 | ||
| RfaH benchmark A multistate protein sequence-design benchmark using the fold-switching conformations of RfaH. | dataset 2024-07-11 | Fully open | 0 | ||
| scBench Agentic evaluation suite for data-grounded single-cell analysis across diverse sequencing technologies and workflow stages. | agentic-eval 2026-02-09 | Partially open | 0 | ||
| scIB A 13-task benchmark of single-cell data integration across simulated, scRNA-seq, and scATAC-seq settings, evaluated with 14 batch-removal and biological-conservation metrics. | suite 2021-12-23 | Fully open | 1 | ||
| SCIGYM An agentic systems-biology suite in which language models iteratively perturb simulated SBML systems, analyze time-series observations in Python, and reconstruct hidden biological reactions. | suite 2025-05-16 | Fully open | 0 | ||
| SCIGYM Large The formally released SCIGYM track containing the 213 systems not included in the creator paper's model evaluation, with systems reaching up to 400 reactions. | track 2025-05-16 | Fully open | 0 | ||
| SCIGYM Small The formally released and creator-evaluated SCIGYM track containing biological systems with fewer than ten reactions. | track 2025-05-16 | Fully open | 2 | ||
| Single-cell Omics Arena A benchmark for evaluating LLM cell-type annotation across scRNA-seq and single-cell multiomics data. | suite 2025-11-24 | Partially open Conflicted · high | 2 | ||
| SpatialBench A benchmark of deterministic, verifiable agentic problems derived from real spatial-transcriptomics workflows, testing whether agents can manipulate data and recover key biological results. | agentic-eval 2025-12-26 | Partially open | 7 | ||
| TAPE A five-task benchmark for protein representation learning spanning secondary structure, residue contacts, remote homology, fluorescence, and stability. | suite 2019-06-19 | Fully open | 1 | ||
| VirBench Retrieval benchmark that tests whether scientific agents can answer verified viral-sequence questions by querying NCBI Virus. | agentic-eval 2025-05-20 | Metadata only | 1 |
Antibody-focused mutational binding data with accompanying structures for computational affinity prediction.
A framework using antibody–antigen complexes to evaluate affinity prediction and antibody redesign.
Private Anthropic computational-biology evaluation direction reported only through a model-trend chart, without public tasks, counts, or protocol details.
An Anthropic private internal suite reported only through an official accuracy chart covering scientific figure interpretation, computational biology, and protein understanding.
Private Anthropic protein-understanding evaluation direction reported only through a model-trend chart, without public tasks, counts, or protocol details.
Private Anthropic evaluation direction for scientific figure interpretation, reported only through a model-trend chart with no task count or released examples.
A living collection of eight curated 3D molecular-learning tasks spanning small molecules, protein interactions and mutations, ligand binding, and protein/RNA structure ranking.
A 13-task RNA representation benchmark covering structure, function, processing, modification, translation, degradation, programmable switches, and CRISPR activity.
A human CRISPR-Cas9 paired-guide library created to compare dual- and single-targeting strategies in loss-of-function screens.
A human CRISPR-Cas9 guide-RNA library assembled to compare single-targeting library performance in loss-of-function screens.
A multi-omics sequence-instruction suite with 21 formal predictive tasks across DNA, RNA, protein, and multi-molecule inputs; the final paper reports task-specific held-out evaluations rather than a single aggregate score.
Binary neutralization prediction for an antibody-antigen protein-sequence pair, evaluated with Matthews correlation coefficient.
Regression of alternative-polyadenylation isoform usage from an RNA sequence, reported as squared Pearson correlation under the paper's R2 label.
Binary detection of a core promoter in a short DNA sequence, evaluated with Matthews correlation coefficient.
Regression of CRISPR guide on-target activity from an RNA sequence, evaluated with Spearman rank correlation.
Two-output regression of housekeeping and developmental enhancer activity from a DNA sequence, evaluated with separate Pearson correlations.
Multi-label Enzyme Commission number prediction from a protein sequence, evaluated with the creator's Fmax implementation.
Binary prediction of whether a DNA sequence carries an epigenetic mark, evaluated with Matthews correlation coefficient.
Binary interaction prediction for enhancer and promoter DNA sequences, evaluated with Matthews correlation coefficient.
Regression of protein fluorescence from an amino-acid sequence, evaluated with Spearman rank correlation.
Multi-label prediction of RNA chemical modifications, evaluated with macro area under the ROC curve.
Regression of mean ribosome loading from an RNA sequence, reported as squared Pearson correlation under the paper's R2 label.
Thirteen-class functional classification of a non-coding RNA sequence, evaluated with exact extracted-label accuracy.
Binary promoter detection in a 300-base-pair DNA context, evaluated with Matthews correlation coefficient.
Three-output regression of ON, OFF, and ON/OFF programmable RNA-switch values, aggregated as their mean squared Pearson correlation.
Binary interaction prediction for an RNA and protein sequence pair, evaluated with Matthews correlation coefficient.
Regression of siRNA efficiency from paired sequence context, evaluated with the SAIS-inspired mixed score.
Binary protein-solubility prediction from an amino-acid sequence, evaluated with accuracy.
Regression of protein stability from an amino-acid sequence, evaluated with Spearman rank correlation.
Binary detection of transcription-factor binding sites in human DNA sequences, evaluated with Matthews correlation coefficient.
Binary detection of transcription-factor binding sites in mouse DNA sequences, evaluated with Matthews correlation coefficient.
Regression of protein thermostability from an amino-acid sequence, evaluated with Spearman rank correlation.
An agentic bioinformatics benchmark of objective, expert-authored mysteries over anonymized real-world biological data, scored on final answers rather than prescribed analysis paths.
Agentic evaluation of pathogen genomic-surveillance workflow selection and analysis from raw or near-raw sequencing data.
A containerized benchmark of long-horizon bioinformatics analysis over real published notebooks and associated data, with open-answer and multiple-choice evaluation modes.
A cross-domain suite for discerning defensible analysis decisions and generating executable end-to-end analyses for open-ended scientific research questions; four of its twelve source questions are explicitly biological or ecological.
The BLADE track requiring a conceptual-variable specification, executable data-transformation function, and statistical-model function for each open-ended research question and dataset.
The BLADE track for selecting the most or least justifiable conceptual-variable and data-transformation decisions for a research question and dataset.
A multistate protein sequence-design benchmark spanning CaM conformations and binding modes.
Weekly, automated, independent, blind evaluation of registered macromolecular structure-prediction servers on complete PDB entries whose experimental structures are withheld during prediction.
Biennial blind community experiments that assess macromolecular structure, complex, ligand, and model-accuracy prediction against experimental structures withheld during prediction.
Dedicated CASP17 category for blind prediction of antibody-antigen, nanobody-antigen, and T-cell receptor complex structures.
Formal CASP track for blind prediction of protein-ligand binding poses, binding affinity or rank, binding pockets, and pose confidence.
Formal CASP track assessing blind predictions of single-protein structures and post hoc protein evaluation units against withheld experimental coordinates.
Formal CASP track assessing blind protein-complex predictions, including overall folds, interfaces, stoichiometry-free phases, and model-selection phases.
A 100-task agent benchmark of objectively gradable computational-biology problems requiring multi-step reasoning, bespoke code, tools, and real-world external resources.
A reusable benchmark for comparing differential transcript usage detection tools across simulated and real transcriptomics data.
Real single-cell RNA-seq data augmented with known gene perturbations for comparing feature-selection methods.
A supervised protein sequence-to-fitness benchmark that turns three experimental landscapes into 15 biologically motivated dataset splits for testing generalization in protein engineering.
Seven supervised splits over sampled and machine-designed AAV2 VP-1 capsid variants, measuring generalization across mutation depth, fitness, and sampled-versus-designed pools.
Five supervised splits over a downsampled, highly epistatic four-site GB1 immunoglobulin-binding landscape, designed to test mutation-depth and low-to-high-fitness generalization.
Three supervised sequence-to-melting-temperature splits spanning all species, human proteins, and a single human cell line, with sequence-cluster-aware train/test separation.
A research-level agent benchmark of 129 synthetic, multistage computational-biology analyses that require iterative QC, statistical modeling, diagnostics, and decision-relevant judgment.
A versioned collection of nine DNA sequence-classification datasets covering regulatory elements, promoters, enhancers, open chromatin, species, and coding-context discrimination.
A reproducible benchmark for de novo molecular design with five distribution-learning tests and twenty goal-directed generation problems in the current v2 suite.
A practical biology-research suite of 2,457 multiple-choice questions across eight broad categories and 31 versioned task files, with public and private contamination-monitoring splits.
Human-hard, multi-step multiple-choice scenarios involving plasmids, DNA fragments, enzymes, and molecular-cloning workflows.
Database-retrieval category spanning 10 genomics, clinical, protein, regulatory, vaccine-response, and viral-PPI tasks.
Identifies genes associated with a phenotype in DisGeNET but not OMIM.
Retrieves human-gene cytogenetic locations from the stated Ensembl release.
Retrieves computationally predicted human miRNA targets from miRDB.
Retrieves genes in Mammalian Phenotype Tumor Ontology gene sets.
Retrieves membership in MSigDB C6 oncogenic-signature gene sets.
Retrieves promoter-region transcription-factor binding-site annotations from GTRD.
Uses a protein sequence and ClinVar lookup to identify benign or pathogenic variants.
Identifies ClinVar variant pathogenicity while reasoning across multiple protein sequences.
Retrieves membership in MSigDB vaccine-response gene sets.
Retrieves predicted human interaction partners of viral proteins from P-HIPSter.
Multiple-choice interpretation and multi-element reasoning over scientific figures shown without captions or paper context.
Literature-retrieval questions whose answers require findings in full research papers rather than titles or abstracts.
Troubleshoots intentionally modified published biological protocols by selecting steps that would repair the stated outcome.
Sequence-comprehension and manipulation category spanning 15 formal tasks involving PCR, restriction digestion, ORFs, translation, GC content, and DNA–protein relationships.
Finds the amino acid encoded at a specified position in the longest ORF of a DNA sequence.
Translates the longest ORF in a DNA sequence to its amino-acid sequence.
Counts open reading frames encoding proteins above a specified amino-acid length.
Selects an RNA sequence whose ORF context is most likely to yield high translation efficiency.
Selects restriction-cloning primers from a named gene and enzyme pair.
Selects primers for Gibson assembly into a HindIII-linearized vector.
Selects primers for Gibson assembly into a SmaI-linearized vector.
Infers restriction enzymes from a gene name and primer pair.
Selects primers that produce a requested amplicon length from a DNA template.
Calculates expected amplicon length from a primer pair and DNA template.
Selects restriction-cloning primers from an explicit gene sequence and enzyme pair.
Selects primers that produce a requested amplicon sequence from a DNA template.
Calculates the rounded GC percentage of a DNA sequence.
Calculates fragment lengths after restriction digestion of a DNA sequence.
Calculates the number of fragments after restriction digestion of a DNA sequence.
Retrieval and interpretation questions answerable from paper supplementary text or PDF tables.
Lookup, calculation, and reasoning questions over table images extracted from scientific papers.
Expert-authored, artifact-rich free-response tasks that evaluate realistic research judgment across applied life-science workflows.
The original molecular-machine-learning benchmark of 17 dataset collections and more than 800 prediction endpoints spanning quantum, physicochemical, biophysical, and physiological properties.
A multistate protein sequence-design benchmark targeting the multispecific PapD binding interface.
A reusable protein-protein binding-affinity dataset with complex structures, measured affinities, receptor and ligand chains, and mutation annotations.
Versioned deep-mutational-scanning and clinical-variant benchmarks for protein fitness prediction and design in zero-shot and supervised regimes.
ProteinGym track for classifying short human clinical insertion and deletion variants against ClinVar and gnomAD-derived labels.
ProteinGym track for classifying expert-annotated human clinical substitution variants on a per-protein basis.
ProteinGym track for predicting experimental fitness measurements of insertion and deletion mutants across deep-mutational-scanning assays.
ProteinGym track for predicting experimental fitness measurements of substitution mutants across deep-mutational-scanning assays.
A creator-curated set of 944 protein-science multiple-choice questions with answer explanations, generated from research literature and released for evaluating text LLM protein understanding.
A multistate protein sequence-design benchmark using the fold-switching conformations of RfaH.
Agentic evaluation suite for data-grounded single-cell analysis across diverse sequencing technologies and workflow stages.
A 13-task benchmark of single-cell data integration across simulated, scRNA-seq, and scATAC-seq settings, evaluated with 14 batch-removal and biological-conservation metrics.
An agentic systems-biology suite in which language models iteratively perturb simulated SBML systems, analyze time-series observations in Python, and reconstruct hidden biological reactions.
The formally released SCIGYM track containing the 213 systems not included in the creator paper's model evaluation, with systems reaching up to 400 reactions.
The formally released and creator-evaluated SCIGYM track containing biological systems with fewer than ten reactions.
A benchmark for evaluating LLM cell-type annotation across scRNA-seq and single-cell multiomics data.
A benchmark of deterministic, verifiable agentic problems derived from real spatial-transcriptomics workflows, testing whether agents can manipulate data and recover key biological results.
A five-task benchmark for protein representation learning spanning secondary structure, residue contacts, remote homology, fluorescence, and stability.
Retrieval benchmark that tests whether scientific agents can answer verified viral-sequence questions by querying NCBI Virus.
No benchmark matches this combination. The URL still preserves your filters.