Evaluation run
genomic-benchmarks-creator-full
From Genomic benchmarks: a collection of datasets for genomic sequence classification
Scopefull · n=9
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
BrowserNot applicable
InternetNot applicable
DatabasesNot applicable
Code executionNot applicable
ContainerNot reported
External toolsTensorFlow and PyTorch CNN baseline workflows
Token budgetNot applicable
Time / cost budgetNot reported
TemperatureNot applicable
SeedNot reported
RepeatsNot reported
Graderdeterministic classification scorer · human review: no
StatisticsAccuracy and F1 are reported independently for each dataset and framework.
ContaminationDataset-specific train/test splits; duplicate and background-generation controls follow each construction notebook.
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | percent | held-out test sequences per dataset | Not reported |
| F1 score | absolute | percent | held-out test sequences per dataset | Not reported |
No numeric result rows are published yet; the verified protocol remains useful.
Evidence
- table: Methods and Tables 1-2 (Lists all nine datasets, train/test design, CNN workflows, and accuracy/F1 results.) — supports /scope, /protocol, /metrics