Evaluation run
lab-bench-dbqa-vax-response-creator-mcq
From LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | proportion | correct / all questions; mean over three runs | exact choice |
| Precision (selective accuracy) | absolute | proportion | correct / attempted questions; mean over three runs | insufficient-information responses excluded from denominator |
| Coverage | absolute | proportion | attempted / all questions; mean over three runs | insufficient-information responses treated as not attempted |
Results
| Model | Metric | Value | n |
|---|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Accuracy | 0.21 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Precision (selective accuracy) | 0.57 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Coverage | 0.36 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Accuracy | 0.14 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Precision (selective accuracy) | 0.65 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Coverage | 0.21 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Accuracy | 0.13 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Precision (selective accuracy) | 0.61 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Coverage | 0.21 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Accuracy | 0.33 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.45 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Coverage | 0.73 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Accuracy | 0.06 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.65 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Coverage | 0.09 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Accuracy | 0.32 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Precision (selective accuracy) | 0.38 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Coverage | 0.83 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Accuracy | 0.15 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Precision (selective accuracy) | 0.54 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Coverage | 0.29 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
Evidence
- section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
- table: Tables 2–4, row DbQA_vax_response_task-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results