Evaluation run
bioinstruction-aan-closed-baselines
Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| MCC | absolute | percent | Matthews correlation coefficient across held-out test examples, scaled by 100 | Not reported |
Results
| Model | Metric | Value | n |
|---|---|---|---|
| GPT-4o-mini (Biology-Instructions snapshot not reported) | MCC | 1.59 percent AAN creator-paper result; metric scaled by 100. | 3301 |
| GPT-4o (Biology-Instructions snapshot not reported) | MCC | -3.29 percent AAN creator-paper result; metric scaled by 100. | 3301 |
Evidence
- table: Table 2 (AAN test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
- table: Table 7 (AAN column) (All registered model values; literature-SOTA row omitted.) — supports /results