Evaluation run
bioinstruction-ncrna-closed-baselines
Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | percent | Exact extracted-class accuracy across held-out test examples, scaled by 100 | Not reported |
Results
| Model | Metric | Value | n |
|---|---|---|---|
| GPT-4o-mini (Biology-Instructions snapshot not reported) | Accuracy | 3 percent ncRNA creator-paper result; metric scaled by 100. | 4840 |
| GPT-4o (Biology-Instructions snapshot not reported) | Accuracy | 5.6 percent ncRNA creator-paper result; metric scaled by 100. | 4840 |
Evidence
- table: Table 2 (ncRNA test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
- table: Table 5 (ncRNA column) (All registered model values; literature-SOTA row omitted.) — supports /results