BEACON
A 13-task RNA representation benchmark covering structure, function, processing, modification, translation, degradation, programmable switches, and CRISPR activity.
Scientific domain
Bulk and other transcriptome-level measurements and analyses.
转录组学
A 13-task RNA representation benchmark covering structure, function, processing, modification, translation, degradation, programmable switches, and CRISPR activity.
A multi-omics sequence-instruction suite with 21 formal predictive tasks across DNA, RNA, protein, and multi-molecule inputs; the final paper reports task-specific held-out evaluations rather than a single aggregate score.
Regression of alternative-polyadenylation isoform usage from an RNA sequence, reported as squared Pearson correlation under the paper's R2 label.
Regression of CRISPR guide on-target activity from an RNA sequence, evaluated with Spearman rank correlation.
Multi-label prediction of RNA chemical modifications, evaluated with macro area under the ROC curve.
Regression of mean ribosome loading from an RNA sequence, reported as squared Pearson correlation under the paper's R2 label.
Thirteen-class functional classification of a non-coding RNA sequence, evaluated with exact extracted-label accuracy.
Three-output regression of ON, OFF, and ON/OFF programmable RNA-switch values, aggregated as their mean squared Pearson correlation.
Binary interaction prediction for an RNA and protein sequence pair, evaluated with Matthews correlation coefficient.
Regression of siRNA efficiency from paired sequence context, evaluated with the SAIS-inspired mixed score.
An agentic bioinformatics benchmark of objective, expert-authored mysteries over anonymized real-world biological data, scored on final answers rather than prescribed analysis paths.
A containerized benchmark of long-horizon bioinformatics analysis over real published notebooks and associated data, with open-answer and multiple-choice evaluation modes.
A 100-task agent benchmark of objectively gradable computational-biology problems requiring multi-step reasoning, bespoke code, tools, and real-world external resources.
A reusable benchmark for comparing differential transcript usage detection tools across simulated and real transcriptomics data.
Real single-cell RNA-seq data augmented with known gene perturbations for comparing feature-selection methods.
A research-level agent benchmark of 129 synthetic, multistage computational-biology analyses that require iterative QC, statistical modeling, diagnostics, and decision-relevant judgment.
A practical biology-research suite of 2,457 multiple-choice questions across eight broad categories and 31 versioned task files, with public and private contamination-monitoring splits.
Database-retrieval category spanning 10 genomics, clinical, protein, regulatory, vaccine-response, and viral-PPI tasks.
Retrieves computationally predicted human miRNA targets from miRDB.
Retrieves membership in MSigDB vaccine-response gene sets.
Sequence-comprehension and manipulation category spanning 15 formal tasks involving PCR, restriction digestion, ORFs, translation, GC content, and DNA–protein relationships.
Selects an RNA sequence whose ORF context is most likely to yield high translation efficiency.
Expert-authored, artifact-rich free-response tasks that evaluate realistic research judgment across applied life-science workflows.
Agentic evaluation suite for data-grounded single-cell analysis across diverse sequencing technologies and workflow stages.
A 13-task benchmark of single-cell data integration across simulated, scRNA-seq, and scATAC-seq settings, evaluated with 14 batch-removal and biological-conservation metrics.
A benchmark for evaluating LLM cell-type annotation across scRNA-seq and single-cell multiomics data.
A benchmark of deterministic, verifiable agentic problems derived from real spatial-transcriptomics workflows, testing whether agents can manipulate data and recover key biological results.
Registry records tagged Transcriptomics, counted by capability.
| Capability | Records |
|---|---|
| Knowledge | 4 |
| Evidence synthesis | 2 |
| Retrieval | 6 |
| Prediction | 15 |
| Classification | 8 |
| Regression | 7 |
| Design | 3 |
| Generation | 1 |
| Optimization | 1 |
| Data analysis | 12 |
| Coding | 6 |
| Tool use | 8 |
| Experiment planning | 2 |
| Troubleshooting | 2 |
| Scientific reasoning | 13 |
| Scientific communication | 1 |
Task mappings are evidence-backed and may be partial for mixed suites.