Evaluation run
bioinstruction-sirna-closed-baselines
Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Mixed Score | absolute | percent | Creator mixed score combining capped MAE and range-MAE-weighted binary F1, scaled by 100 | Not reported |
Results
| Model | Metric | Value | n |
|---|---|---|---|
| GPT-4o-mini (Biology-Instructions snapshot not reported) | Mixed Score | 30.37 percent siRNA creator-paper result; metric scaled by 100. | 6688 |
| GPT-4o (Biology-Instructions snapshot not reported) | Mixed Score | 0 percent siRNA creator-paper result; metric scaled by 100. | 6688 |
Evidence
- table: Table 2 (siRNA test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
- table: Table 7 (siRNA column) (All registered model values; literature-SOTA row omitted.) — supports /results