Scientific domain

Transcriptomics

Bulk and other transcriptome-level measurements and analyses.

转录组学

27 registered benchmarks and tracks

suite1 runs

BEACON

A 13-task RNA representation benchmark covering structure, function, processing, modification, translation, degradation, programmable switches, and CRISPR activity.

PredictionClassificationRegression
suite0 runs

Biology-Instructions

A multi-omics sequence-instruction suite with 21 formal predictive tasks across DNA, RNA, protein, and multi-molecule inputs; the final paper reports task-specific held-out evaluations rather than a single aggregate score.

PredictionClassificationRegressionScientific reasoning
track3 runs

Biology-Instructions APA Isoform Prediction

Regression of alternative-polyadenylation isoform usage from an RNA sequence, reported as squared Pearson correlation under the paper's R2 label.

PredictionRegression
agentic-eval3 runs

BioMysteryBench

An agentic bioinformatics benchmark of objective, expert-authored mysteries over anonymized real-world biological data, scored on final answers rather than prescribed analysis paths.

Data analysisCodingTool useScientific reasoning
agentic-eval11 runs

BixBench

A containerized benchmark of long-horizon bioinformatics analysis over real published notebooks and associated data, with open-answer and multiple-choice evaluation modes.

Data analysisCodingTool useScientific reasoning
agentic-eval9 runs

CompBioBench

A 100-task agent benchmark of objectively gradable computational-biology problems requiring multi-step reasoning, bespoke code, tools, and real-world external resources.

Data analysisCodingTool useRetrieval
suite0 runs

crafted experiments

Real single-cell RNA-seq data augmented with known gene perturbations for comparing feature-selection methods.

Data analysis
agentic-eval13 runs

GeneBench-Pro

A research-level agent benchmark of 129 synthetic, multistage computational-biology analyses that require iterative QC, statistical modeling, diagnostics, and decision-relevant judgment.

Data analysisCodingTool useScientific reasoning
suite0 runs

LAB-Bench

A practical biology-research suite of 2,457 multiple-choice questions across eight broad categories and 31 versioned task files, with public and private contamination-monitoring splits.

KnowledgeEvidence synthesisRetrievalPrediction
track0 runs

LAB-Bench DbQA

Database-retrieval category spanning 10 genomics, clinical, protein, regulatory, vaccine-response, and viral-PPI tasks.

RetrievalKnowledgePredictionClassification
track1 runs

LAB-Bench SeqQA

Sequence-comprehension and manipulation category spanning 15 formal tasks involving PCR, restriction digestion, ORFs, translation, GC content, and DNA–protein relationships.

Data analysisDesignPredictionScientific reasoning
agentic-eval1 runs

LifeSciBench

Expert-authored, artifact-rich free-response tasks that evaluate realistic research judgment across applied life-science workflows.

Evidence synthesisRetrievalDesignGeneration
agentic-eval0 runs

scBench

Agentic evaluation suite for data-grounded single-cell analysis across diverse sequencing technologies and workflow stages.

Data analysisCodingTool useScientific reasoning
suite1 runs

scIB

A 13-task benchmark of single-cell data integration across simulated, scRNA-seq, and scATAC-seq settings, evaluated with 14 batch-removal and biological-conservation metrics.

Data analysis
suite2 runs

Single-cell Omics Arena

A benchmark for evaluating LLM cell-type annotation across scRNA-seq and single-cell multiomics data.

ClassificationScientific reasoning
agentic-eval7 runs

SpatialBench

A benchmark of deterministic, verifiable agentic problems derived from real spatial-transcriptomics workflows, testing whether agents can manipulate data and recover key biological results.

Data analysisCodingTool useScientific reasoning

Capability coverage

Registry records tagged Transcriptomics, counted by capability.

CSV ↓
Accessible data table
CapabilityRecords
Knowledge4
Evidence synthesis2
Retrieval6
Prediction15
Classification8
Regression7
Design3
Generation1
Optimization1
Data analysis12
Coding6
Tool use8
Experiment planning2
Troubleshooting2
Scientific reasoning13
Scientific communication1
Coverage-gap reading: zero counts indicate a gap in this registry, not proof that no benchmark exists. Propose a primary source through the contribution form.