Evaluation run
beacon-creator-full
From BEACON: Benchmark for Comprehensive RNA Tasks and Language Models
Scopefull · n=13
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
BrowserNot applicable
InternetNot applicable
DatabasesNot applicable
Code executionNot applicable
ContainerNot reported
External toolstask-specific fine-tuning scripts over CNN/ResNet/LSTM and RNA language models
Token budgetNot applicable
Time / cost budgetNot reported
TemperatureNot applicable
SeedNot reported
Repeats3
Graderdeterministic task-specific scorer · human review: no
StatisticsTask-level mean and variation over three seeds; no cross-task total score.
ContaminationTask-specific source datasets and prescribed train/validation/test partitions.
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| F1 | absolute | percent | task-specific | Not reported |
| Top L Precision | absolute | percent | contact-map task | Not reported |
| R2 | absolute | percent | task-specific regression | Not reported |
| Top-k ACC | absolute | percent | splice-site task | Not reported |
| ACC | absolute | percent | sequence classification | Not reported |
| AUC | absolute | percent | modification prediction | Not reported |
| MCRMSE | absolute | error | mean columnwise RMSE | Not reported |
| Spearman Corr | absolute | percent correlation | CRISPR task examples | Not reported |
No numeric result rows are published yet; the verified protocol remains useful.
Evidence
- table: Tables 1 and 3; Sections 5.1-5.2; Appendix A.1 (Gives split sizes, task metrics, model families, three random seeds, and task-specific training settings.) — supports /scope, /protocol, /metrics