Evaluation run
bioinstruction-rpi-closed-baselines
Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| MCC | absolute | percent | Matthews correlation coefficient across held-out test examples, scaled by 100 | Not reported |
Results
| Model | Metric | Value | n |
|---|---|---|---|
| GPT-4o-mini (Biology-Instructions snapshot not reported) | MCC | 1.22 percent RPI creator-paper result; metric scaled by 100. | 4164 |
| GPT-4o (Biology-Instructions snapshot not reported) | MCC | 1.17 percent RPI creator-paper result; metric scaled by 100. | 4164 |
Evidence
- table: Table 2 (RPI test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
- table: Table 7 (RPI column) (All registered model values; literature-SOTA row omitted.) — supports /results