paper · benchmark creator

BEACON: Benchmark for Comprehensive RNA Tasks and Language Models

Shanghai Artificial Intelligence Laboratory · University of Sydney · University of Hong Kong · Shanghai Jiao Tong University · Fudan University · Chinese University of Hong Kong · 2024-12-10

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

No BenchmarkUse relation is normalized for this legacy work yet. Existing EvaluationRuns remain available below.

Normalized evaluation runs

BEACON1 run

Open benchmark record →

beacon-creator-task-nativevneurips-2024
Scopefull · n=13
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
BrowserNot applicable
InternetNot applicable
DatabasesNot applicable
Code executionNot applicable
ContainerNot reported
External toolstask-specific fine-tuning scripts over CNN/ResNet/LSTM and RNA language models
Token budgetNot applicable
Time / cost budgetNot reported
TemperatureNot applicable
SeedNot reported
Repeats3
Graderdeterministic task-specific scorer · human review: no
StatisticsTask-level mean and variation over three seeds; no cross-task total score.
ContaminationTask-specific source datasets and prescribed train/validation/test partitions.
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
F1absolutepercenttask-specificNot reported
Top L Precisionabsolutepercentcontact-map taskNot reported
R2absolutepercenttask-specific regressionNot reported
Top-k ACCabsolutepercentsplice-site taskNot reported
ACCabsolutepercentsequence classificationNot reported
AUCabsolutepercentmodification predictionNot reported
MCRMSEabsoluteerrormean columnwise RMSENot reported
Spearman Corrabsolutepercent correlationCRISPR task examplesNot reported

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • table: Tables 1 and 3; Sections 5.1-5.2; Appendix A.1 (Gives split sizes, task metrics, model families, three random seeds, and task-specific training settings.) — supports /scope, /protocol, /metrics