Evaluation run
lab-bench-seqqa-orf-seq-numlen-creator-mcq
From LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | proportion | correct / all questions; mean over three runs | exact choice |
| Precision (selective accuracy) | absolute | proportion | correct / attempted questions; mean over three runs | insufficient-information responses excluded from denominator |
| Coverage | absolute | proportion | attempted / all questions; mean over three runs | insufficient-information responses treated as not attempted |
Results
| Model | Metric | Value | n |
|---|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Accuracy | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Precision (selective accuracy) | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Coverage | 0.01 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Accuracy | 0.07 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Precision (selective accuracy) | 0.28 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Coverage | 0.26 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Accuracy | 0.05 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Precision (selective accuracy) | 0.18 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Coverage | 0.31 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Accuracy | 0.11 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.14 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Coverage | 0.81 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Accuracy | 0.07 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.27 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Coverage | 0.29 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Accuracy | 0.26 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Precision (selective accuracy) | 0.26 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Coverage | 1 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Accuracy | 0.23 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Precision (selective accuracy) | 0.29 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Coverage | 0.81 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
Evidence
- section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
- table: Tables 2–4, row SeqQA_ORF-seq-numlen-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results