Evaluation run
bioinstruction-solubility-closed-baselines
Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | percent | Exact binary accuracy across held-out test examples, scaled by 100 | Not reported |
Results
| Model | Metric | Value | n |
|---|---|---|---|
| GPT-4o-mini (Biology-Instructions snapshot not reported) | Accuracy | 50.02 percent Sol creator-paper result; metric scaled by 100. | 2001 |
| GPT-4o (Biology-Instructions snapshot not reported) | Accuracy | 51.67 percent Sol creator-paper result; metric scaled by 100. | 2001 |
Evidence
- table: Table 2 (Sol test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
- table: Table 6 (Sol column) (All registered model values; literature-SOTA row omitted.) — supports /results