LAB-Bench CloningScenarios
Human-hard, multi-step multiple-choice scenarios involving plasmids, DNA fragments, enzymes, and molecular-cloning workflows.
4 evaluation run(s)
suite · audited-with-caveats · verified 2026-07-21
A practical biology-research suite of 2,457 multiple-choice questions across eight broad categories and 31 versioned task files, with public and private contamination-monitoring splits.
Benchmark definition
| Version | Status | Release / as-of | Total | Formal tracks |
|---|---|---|---|---|
paper-v3lab-bench-paper-v3 | active | 2024-07-17 | 2457 (multiple-choice questions across the complete public and private creator snapshot) | lab-bench-litqa2, lab-bench-suppqa, lab-bench-figqa, lab-bench-tableqa, lab-bench-dbqa, lab-bench-protocolqa, lab-bench-seqqa, lab-bench-cloning-scenarios |
repository-998a8e0lab-bench-repository-998a8e0 | current | 2025-09-27 | 2457 (multiple-choice questions across the complete public and private creator snapshot) | lab-bench-litqa2, lab-bench-suppqa, lab-bench-figqa, lab-bench-tableqa, lab-bench-dbqa, lab-bench-protocolqa, lab-bench-seqqa, lab-bench-cloning-scenarios |
| ID | Count | Basis | Partition? | Notes |
|---|---|---|---|---|
LitQA2lab-bench-litqa2 | 248 | questions | Exclusive & exhaustive | — |
SuppQAlab-bench-suppqa | 102 | questions | Exclusive & exhaustive | — |
FigQAlab-bench-figqa | 226 | questions | Exclusive & exhaustive | — |
TableQAlab-bench-tableqa | 305 | questions | Exclusive & exhaustive | — |
DbQAlab-bench-dbqa | 650 | questions | Exclusive & exhaustive | — |
ProtocolQAlab-bench-protocolqa | 135 | questions | Exclusive & exhaustive | — |
SeqQAlab-bench-seqqa | 750 | questions | Exclusive & exhaustive | — |
CloningScenarioslab-bench-cloning-scenarios | 41 | questions | Exclusive & exhaustive | — |
Broad categorieslab-bench-broad-categories | 8 | categories | No | Not a question count. |
Versioned formal task fileslab-bench-formal-task-files | 31 | versioned split files | No | The 31 files are the six single-task categories, 10 DbQA subtasks, 15 SeqQA subtasks, and CloningScenarios. |
Narrower subtasks stated in READMElab-bench-readme-narrower-subtasks | 30Conflicted · high | subtasks claimed in README | No | Conflicts with the 31 versioned split files and 31 creator-result rows. |
Human-hard, multi-step multiple-choice scenarios involving plasmids, DNA fragments, enzymes, and molecular-cloning workflows.
4 evaluation run(s)
Database-retrieval category spanning 10 genomics, clinical, protein, regulatory, vaccine-response, and viral-PPI tasks.
0 evaluation run(s)
Multiple-choice interpretation and multi-element reasoning over scientific figures shown without captions or paper context.
5 evaluation run(s)
Literature-retrieval questions whose answers require findings in full research papers rather than titles or abstracts.
1 evaluation run(s)
Troubleshoots intentionally modified published biological protocols by selecting steps that would repair the stated outcome.
4 evaluation run(s)
Sequence-comprehension and manipulation category spanning 15 formal tasks involving PCR, restriction digestion, ORFs, translation, GC content, and DNA–protein relationships.
1 evaluation run(s)
Retrieval and interpretation questions answerable from paper supplementary text or PDF tables.
1 evaluation run(s)
Lookup, calculation, and reasoning questions over table images extracted from scientific papers.
1 evaluation run(s)
Scientific Task Atlas
partial for repository-998a8e0. Formal child tracks support several precise mappings; the mixed suite has no exhaustive creator scientific-task taxonomy.
| Scientific task | Coverage | Count | Mapping | Evidence |
|---|---|---|---|---|
| Scientific database retrieval | explicitly-in-scope | 650 questions Questions across the ten formal DbQA child tasks. | official-track high confidence | lab-bench-evidence-repositoryCount is specific to DbQA and is not added to other task claims. |
| Scientific evidence interpretation | explicitly-in-scope | Not reported FigQA, LitQA2, SuppQA, and TableQA questions. | official-track high confidence | lab-bench-evidence-paperThe root record does not publish this cross-track subtotal. |
| Experiment and protocol planning | explicitly-in-scope | 135 questions ProtocolQA questions across public and private splits. | official-track high confidence | lab-bench-evidence-paperCloningScenarios is shown separately and is not included in this count. |
| Protein-protein interaction prediction | explicitly-in-scope | 50 questions Viral PPI formal-task questions. | official-track high confidence | lab-bench-evidence-paperViral-human PPI database-retrieval questions. |
| Protein sequence design | not-in-scope | 0 questions Released LAB-Bench questions. | official-taxonomy high confidence | lab-bench-evidence-paperPrimer and cloning design are not relabeled as protein sequence design. |
| Domain | Coverage | Count | Interpretation |
|---|---|---|---|
| Protein sequence | explicitly-in-scope | Not reported | SeqQA and two ClinVar DbQA tasks require DNA/protein-sequence reasoning, but the sources do not publish a protein-only question count. |
| Protein-protein binding | explicitly-in-scope | 50 | The Viral PPI formal task has 50 questions about predicted viral–human protein interactions. |
| Protein design | not-in-scope | 0 | Primer and cloning design are in scope; protein sequence design is not a released LAB-Bench task. |
Evaluation registry
A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.
Evaluation run
Evaluated models / systems: Claude Opus 4, Claude Opus 4.1, Claude Sonnet 4, Claude Sonnet 4.5
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| LAB-Bench score | absolute | proportion | Not reported | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude Sonnet 4 | LAB-Bench score | 0.485 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
| Claude Opus 4 | LAB-Bench score | 0.545 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
| Claude Opus 4.1 | LAB-Bench score | 0.758 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
| Claude Sonnet 4.5 | LAB-Bench score | 0.667 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
Evaluation run
From LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | proportion | correct / all questions; mean over three runs | exact choice |
| Precision (selective accuracy) | absolute | proportion | correct / attempted questions; mean over three runs | insufficient-information responses excluded from denominator |
| Coverage | absolute | proportion | attempted / all questions; mean over three runs | insufficient-information responses treated as not attempted |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Accuracy | 0.28 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 41 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Precision (selective accuracy) | 0.54 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 41 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Coverage | 0.52 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 41 |
| Claude 3 Opus (claude-3-opus-20240229) | Accuracy | 0.27 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 41 |
| Claude 3 Opus (claude-3-opus-20240229) | Precision (selective accuracy) | 0.41 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 41 |
| Claude 3 Opus (claude-3-opus-20240229) | Coverage | 0.65 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 41 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Accuracy | 0.1 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 41 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Precision (selective accuracy) | 0.33 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 41 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Coverage | 0.3 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 41 |
| GPT-4o (LAB-Bench snapshot not reported) | Accuracy | 0.28 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 41 |
| GPT-4o (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.37 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 41 |
| GPT-4o (LAB-Bench snapshot not reported) | Coverage | 0.77 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 41 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Accuracy | 0.09 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 41 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.36 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 41 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Coverage | 0.31 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 41 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Accuracy | 0.26 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 41 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Precision (selective accuracy) | 0.29 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 41 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Coverage | 0.9 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 41 |
Evaluation run
From LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Evaluated models / systems: Meta-Llama-3-70B-Instruct (Anyscale API)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | proportion | correct / all questions; mean over three runs | exact choice |
| Precision (selective accuracy) | absolute | proportion | correct / attempted questions; mean over three runs | insufficient-information responses excluded from denominator |
| Coverage | absolute | proportion | attempted / all questions; mean over three runs | insufficient-information responses treated as not attempted |
| Model | Metric | Value | n |
|---|---|---|---|
| Meta-Llama-3-70B-Instruct (Anyscale API) | Accuracy | 0.24 proportion Only 25 prompts fit; 16 were counted as insufficient-information responses for accuracy and coverage. | 41 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Precision (selective accuracy) | 0.34 proportion Only 25 prompts fit; 16 were counted as insufficient-information responses for accuracy and coverage. | 41 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Coverage | 0.72 proportion Only 25 prompts fit; 16 were counted as insufficient-information responses for accuracy and coverage. | 41 |
Evaluation run
From LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), GPT-4o (LAB-Bench snapshot not reported)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | proportion | correct / sampled modified questions | expert judgment against ideal multiple-choice answer |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Accuracy | 0.2 proportion Creator Table 5 open-response study. | 10 |
| GPT-4o (LAB-Bench snapshot not reported) | Accuracy | 0.2 proportion Creator Table 5 open-response study. | 10 |
Evaluated models / systems: Claude Opus 4, Claude Opus 4.1, Claude Sonnet 4, Claude Sonnet 4.5
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| LAB-Bench score | absolute | proportion | Not reported | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude Sonnet 4 | LAB-Bench score | 0.398 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
| Claude Opus 4 | LAB-Bench score | 0.508 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
| Claude Opus 4.1 | LAB-Bench score | 0.481 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
| Claude Sonnet 4.5 | LAB-Bench score | 0.497 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
Evaluation run
From LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | proportion | correct / all questions; mean over three runs | exact choice |
| Precision (selective accuracy) | absolute | proportion | correct / attempted questions; mean over three runs | insufficient-information responses excluded from denominator |
| Coverage | absolute | proportion | attempted / all questions; mean over three runs | insufficient-information responses treated as not attempted |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Accuracy | 0.46 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 226 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Precision (selective accuracy) | 0.54 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 226 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Coverage | 0.85 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 226 |
| Claude 3 Opus (claude-3-opus-20240229) | Accuracy | 0.24 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 226 |
| Claude 3 Opus (claude-3-opus-20240229) | Precision (selective accuracy) | 0.31 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 226 |
| Claude 3 Opus (claude-3-opus-20240229) | Coverage | 0.78 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 226 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Accuracy | 0.25 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 226 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Precision (selective accuracy) | 0.34 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 226 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Coverage | 0.73 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 226 |
| GPT-4o (LAB-Bench snapshot not reported) | Accuracy | 0.29 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 226 |
| GPT-4o (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.3 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 226 |
| GPT-4o (LAB-Bench snapshot not reported) | Coverage | 0.97 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 226 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Accuracy | 0.25 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 226 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.28 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 226 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Coverage | 0.91 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 226 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Accuracy | 0.23 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 226 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Precision (selective accuracy) | 0.24 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 226 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Coverage | 0.95 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 226 |
Evaluation run
From LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), GPT-4o (LAB-Bench snapshot not reported)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | proportion | correct / sampled modified questions | expert judgment against ideal multiple-choice answer |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Accuracy | 0.3 proportion Creator Table 5 open-response study. | 10 |
| GPT-4o (LAB-Bench snapshot not reported) | Accuracy | 0.3 proportion Creator Table 5 open-response study. | 10 |
Evaluated models / systems: Claude Opus 4.6, Claude Sonnet 4.5, Claude Sonnet 4.6
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| FigQA score | absolute | percent | mean over five runs | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude Sonnet 4.6 | FigQA score | 77.1 percent 95% CI is plotted but numeric bounds are not reported. | Not reported |
| Claude Sonnet 4.5 | FigQA score | 59.3 percent 95% CI is plotted but numeric bounds are not reported. | Not reported |
| Claude Opus 4.6 | FigQA score | 78.3 percent 95% CI is plotted but numeric bounds are not reported. | Not reported |
Evaluated models / systems: Claude Opus 4.6, Claude Sonnet 4.5, Claude Sonnet 4.6
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| FigQA score | absolute | percent | mean over five runs | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude Sonnet 4.6 | FigQA score | 58.8 percent 95% CI is plotted but numeric bounds are not reported. | Not reported |
| Claude Sonnet 4.5 | FigQA score | 53.4 percent 95% CI is plotted but numeric bounds are not reported. | Not reported |
| Claude Opus 4.6 | FigQA score | 58 percent 95% CI is plotted but numeric bounds are not reported. | Not reported |
Evaluation run
From LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | proportion | correct / all questions; mean over three runs | exact choice |
| Precision (selective accuracy) | absolute | proportion | correct / attempted questions; mean over three runs | insufficient-information responses excluded from denominator |
| Coverage | absolute | proportion | attempted / all questions; mean over three runs | insufficient-information responses treated as not attempted |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Accuracy | 0.06 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 248 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Precision (selective accuracy) | 0.47 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 248 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Coverage | 0.12 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 248 |
| Claude 3 Opus (claude-3-opus-20240229) | Accuracy | 0.06 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 248 |
| Claude 3 Opus (claude-3-opus-20240229) | Precision (selective accuracy) | 0.43 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 248 |
| Claude 3 Opus (claude-3-opus-20240229) | Coverage | 0.14 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 248 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Accuracy | 0.1 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 248 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Precision (selective accuracy) | 0.44 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 248 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Coverage | 0.23 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 248 |
| GPT-4o (LAB-Bench snapshot not reported) | Accuracy | 0.27 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 248 |
| GPT-4o (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.46 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 248 |
| GPT-4o (LAB-Bench snapshot not reported) | Coverage | 0.58 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 248 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Accuracy | 0.12 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 248 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.4 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 248 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Coverage | 0.31 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 248 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Accuracy | 0.27 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 248 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Precision (selective accuracy) | 0.38 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 248 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Coverage | 0.7 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 248 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Accuracy | 0.35 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 248 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Precision (selective accuracy) | 0.38 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 248 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Coverage | 0.92 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 248 |
Evaluated models / systems: Claude Sonnet 4, Claude Sonnet 4.5
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| ProtocolQA score | absolute | proportion | Not reported | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude Sonnet 4.5 | ProtocolQA score | 0.83 proportion Official release rounds the system-card point label 0.833 to 0.83. | Not reported |
| Claude Sonnet 4 | ProtocolQA score | 0.74 proportion Official release rounds the system-card point label 0.741 to 0.74. | Not reported |
Evaluated models / systems: Claude Opus 4, Claude Opus 4.1
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| LAB-Bench score | absolute | proportion | Not reported | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude Opus 4 | LAB-Bench score | 0.796 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
| Claude Opus 4.1 | LAB-Bench score | 0.833 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
Evaluation run
From LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | proportion | correct / all questions; mean over three runs | exact choice |
| Precision (selective accuracy) | absolute | proportion | correct / attempted questions; mean over three runs | insufficient-information responses excluded from denominator |
| Coverage | absolute | proportion | attempted / all questions; mean over three runs | insufficient-information responses treated as not attempted |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Accuracy | 0.48 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 135 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Precision (selective accuracy) | 0.66 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 135 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Coverage | 0.73 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 135 |
| Claude 3 Opus (claude-3-opus-20240229) | Accuracy | 0.52 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 135 |
| Claude 3 Opus (claude-3-opus-20240229) | Precision (selective accuracy) | 0.62 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 135 |
| Claude 3 Opus (claude-3-opus-20240229) | Coverage | 0.84 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 135 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Accuracy | 0.47 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 135 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Precision (selective accuracy) | 0.58 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 135 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Coverage | 0.81 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 135 |
| GPT-4o (LAB-Bench snapshot not reported) | Accuracy | 0.53 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 135 |
| GPT-4o (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.56 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 135 |
| GPT-4o (LAB-Bench snapshot not reported) | Coverage | 0.95 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 135 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Accuracy | 0.37 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 135 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.59 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 135 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Coverage | 0.62 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 135 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Accuracy | 0.45 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 135 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Precision (selective accuracy) | 0.51 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 135 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Coverage | 0.87 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 135 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Accuracy | 0.44 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 135 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Precision (selective accuracy) | 0.49 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 135 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Coverage | 0.9 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 135 |
Evaluation run
From LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), GPT-4o (LAB-Bench snapshot not reported)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | proportion | correct / sampled modified questions | expert judgment against ideal multiple-choice answer |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Accuracy | 0.3 proportion Creator Table 5 open-response study. | 20 |
| GPT-4o (LAB-Bench snapshot not reported) | Accuracy | 0.2 proportion Creator Table 5 open-response study. | 20 |
Evaluated models / systems: Claude Opus 4, Claude Opus 4.1, Claude Sonnet 4, Claude Sonnet 4.5
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| LAB-Bench score | absolute | proportion | Not reported | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude Sonnet 4 | LAB-Bench score | 0.682 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
| Claude Opus 4 | LAB-Bench score | 0.723 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
| Claude Opus 4.1 | LAB-Bench score | 0.785 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
| Claude Sonnet 4.5 | LAB-Bench score | 0.78 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
Evaluation run
From LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | proportion | correct / all questions; mean over three runs | exact choice |
| Precision (selective accuracy) | absolute | proportion | correct / attempted questions; mean over three runs | insufficient-information responses excluded from denominator |
| Coverage | absolute | proportion | attempted / all questions; mean over three runs | insufficient-information responses treated as not attempted |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Accuracy | 0.02 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 102 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Precision (selective accuracy) | 0.75 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 102 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Coverage | 0.04 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 102 |
| Claude 3 Opus (claude-3-opus-20240229) | Accuracy | 0.02 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 102 |
| Claude 3 Opus (claude-3-opus-20240229) | Precision (selective accuracy) | 0.59 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 102 |
| Claude 3 Opus (claude-3-opus-20240229) | Coverage | 0.04 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 102 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Accuracy | 0.01 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 102 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Precision (selective accuracy) | 0.32 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 102 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Coverage | 0.04 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 102 |
| GPT-4o (LAB-Bench snapshot not reported) | Accuracy | 0.13 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 102 |
| GPT-4o (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.47 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 102 |
| GPT-4o (LAB-Bench snapshot not reported) | Coverage | 0.29 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 102 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Accuracy | 0.06 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 102 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.39 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 102 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Coverage | 0.14 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 102 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Accuracy | 0.18 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 102 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Precision (selective accuracy) | 0.35 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 102 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Coverage | 0.53 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 102 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Accuracy | 0.2 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 102 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Precision (selective accuracy) | 0.27 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 102 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Coverage | 0.74 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 102 |
Evaluation run
From LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | proportion | correct / all questions; mean over three runs | exact choice |
| Precision (selective accuracy) | absolute | proportion | correct / attempted questions; mean over three runs | insufficient-information responses excluded from denominator |
| Coverage | absolute | proportion | attempted / all questions; mean over three runs | insufficient-information responses treated as not attempted |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Accuracy | 0.83 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 305 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Precision (selective accuracy) | 0.9 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 305 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Coverage | 0.92 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 305 |
| Claude 3 Opus (claude-3-opus-20240229) | Accuracy | 0.67 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 305 |
| Claude 3 Opus (claude-3-opus-20240229) | Precision (selective accuracy) | 0.74 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 305 |
| Claude 3 Opus (claude-3-opus-20240229) | Coverage | 0.9 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 305 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Accuracy | 0.59 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 305 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Precision (selective accuracy) | 0.71 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 305 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Coverage | 0.83 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 305 |
| GPT-4o (LAB-Bench snapshot not reported) | Accuracy | 0.71 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 305 |
| GPT-4o (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.75 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 305 |
| GPT-4o (LAB-Bench snapshot not reported) | Coverage | 0.95 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 305 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Accuracy | 0.54 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 305 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.58 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 305 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Coverage | 0.93 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 305 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Accuracy | 0.49 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 305 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Precision (selective accuracy) | 0.51 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 305 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Coverage | 0.96 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 305 |
lab-bench-cloning-scenarios-anthropic-sonnet45-system-card · lab-bench-cloning-scenarios-anthropic-10shot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude Sonnet 4 | 0.485 | lab-bench-cloning-scenarios-anthropic-10shot-no-tools |
| Claude Opus 4 | 0.545 | lab-bench-cloning-scenarios-anthropic-10shot-no-tools |
| Claude Opus 4.1 | 0.758 | lab-bench-cloning-scenarios-anthropic-10shot-no-tools |
| Claude Sonnet 4.5 | 0.667 | lab-bench-cloning-scenarios-anthropic-10shot-no-tools |
lab-bench-cloning-scenarios-creator-mcq · lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.28 | lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.27 | lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.1 | lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.28 | lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.09 | lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.26 | lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools |
lab-bench-cloning-scenarios-creator-mcq · lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.54 | lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.41 | lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.33 | lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.37 | lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.36 | lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.29 | lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools |
lab-bench-cloning-scenarios-creator-mcq · lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.52 | lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.65 | lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.3 | lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.77 | lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.31 | lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.9 | lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools |
lab-bench-cloning-scenarios-creator-mcq-llama-context · lab-bench-cloning-scenarios-paper-v3-llama-context-limited
| Model | Value | Comparability group |
|---|---|---|
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.24 | lab-bench-cloning-scenarios-paper-v3-llama-context-limited |
lab-bench-cloning-scenarios-creator-mcq-llama-context · lab-bench-cloning-scenarios-paper-v3-llama-context-limited
| Model | Value | Comparability group |
|---|---|---|
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.34 | lab-bench-cloning-scenarios-paper-v3-llama-context-limited |
lab-bench-cloning-scenarios-creator-mcq-llama-context · lab-bench-cloning-scenarios-paper-v3-llama-context-limited
| Model | Value | Comparability group |
|---|---|---|
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.72 | lab-bench-cloning-scenarios-paper-v3-llama-context-limited |
lab-bench-cloning-scenarios-creator-open-response · lab-bench-cloning-scenarios-paper-v3-open-response
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.2 | lab-bench-cloning-scenarios-paper-v3-open-response |
| GPT-4o (LAB-Bench snapshot not reported) | 0.2 | lab-bench-cloning-scenarios-paper-v3-open-response |
lab-bench-figqa-anthropic-sonnet45-system-card · lab-bench-figqa-anthropic-10shot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude Sonnet 4 | 0.398 | lab-bench-figqa-anthropic-10shot-no-tools |
| Claude Opus 4 | 0.508 | lab-bench-figqa-anthropic-10shot-no-tools |
| Claude Opus 4.1 | 0.481 | lab-bench-figqa-anthropic-10shot-no-tools |
| Claude Sonnet 4.5 | 0.497 | lab-bench-figqa-anthropic-10shot-no-tools |
lab-bench-figqa-creator-mcq · lab-bench-figqa-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.46 | lab-bench-figqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.24 | lab-bench-figqa-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.25 | lab-bench-figqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.29 | lab-bench-figqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.25 | lab-bench-figqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.23 | lab-bench-figqa-paper-v3-zero-shot-cot-no-tools |
lab-bench-figqa-creator-mcq · lab-bench-figqa-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.54 | lab-bench-figqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.31 | lab-bench-figqa-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.34 | lab-bench-figqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.3 | lab-bench-figqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.28 | lab-bench-figqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.24 | lab-bench-figqa-paper-v3-zero-shot-cot-no-tools |
lab-bench-figqa-creator-mcq · lab-bench-figqa-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.85 | lab-bench-figqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.78 | lab-bench-figqa-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.73 | lab-bench-figqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.97 | lab-bench-figqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.91 | lab-bench-figqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.95 | lab-bench-figqa-paper-v3-zero-shot-cot-no-tools |
lab-bench-figqa-creator-open-response · lab-bench-figqa-paper-v3-open-response
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.3 | lab-bench-figqa-paper-v3-open-response |
| GPT-4o (LAB-Bench snapshot not reported) | 0.3 | lab-bench-figqa-paper-v3-open-response |
lab-bench-figqa-crop-tool · lab-bench-figqa-sonnet46-adaptive-max-crop-tool-five-runs
| Model | Value | Comparability group |
|---|---|---|
| Claude Sonnet 4.6 | 77.1 | lab-bench-figqa-sonnet46-adaptive-max-crop-tool-five-runs |
| Claude Sonnet 4.5 | 59.3 | lab-bench-figqa-sonnet46-adaptive-max-crop-tool-five-runs |
| Claude Opus 4.6 | 78.3 | lab-bench-figqa-sonnet46-adaptive-max-crop-tool-five-runs |
lab-bench-figqa-no-tools · lab-bench-figqa-sonnet46-adaptive-max-no-tools-five-runs
| Model | Value | Comparability group |
|---|---|---|
| Claude Sonnet 4.6 | 58.8 | lab-bench-figqa-sonnet46-adaptive-max-no-tools-five-runs |
| Claude Sonnet 4.5 | 53.4 | lab-bench-figqa-sonnet46-adaptive-max-no-tools-five-runs |
| Claude Opus 4.6 | 58 | lab-bench-figqa-sonnet46-adaptive-max-no-tools-five-runs |
lab-bench-litqa2-creator-mcq · lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.06 | lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.06 | lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.1 | lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.27 | lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.12 | lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.27 | lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.35 | lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools |
lab-bench-litqa2-creator-mcq · lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.47 | lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.43 | lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.44 | lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.46 | lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.4 | lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.38 | lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.38 | lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools |
lab-bench-litqa2-creator-mcq · lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.12 | lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.14 | lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.23 | lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.58 | lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.31 | lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.7 | lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.92 | lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools |
lab-bench-protocolqa-anthropic · lab-bench-protocolqa-anthropic-life-sciences-10shot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude Sonnet 4.5 | 0.83 | lab-bench-protocolqa-anthropic-life-sciences-10shot-no-tools |
| Claude Sonnet 4 | 0.74 | lab-bench-protocolqa-anthropic-life-sciences-10shot-no-tools |
lab-bench-protocolqa-anthropic-sonnet45-system-card · lab-bench-protocolqa-anthropic-10shot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude Opus 4 | 0.796 | lab-bench-protocolqa-anthropic-10shot-no-tools |
| Claude Opus 4.1 | 0.833 | lab-bench-protocolqa-anthropic-10shot-no-tools |
lab-bench-protocolqa-creator-mcq · lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.48 | lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.52 | lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.47 | lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.53 | lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.37 | lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.45 | lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.44 | lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools |
lab-bench-protocolqa-creator-mcq · lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.66 | lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.62 | lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.58 | lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.56 | lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.59 | lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.51 | lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.49 | lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools |
lab-bench-protocolqa-creator-mcq · lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.73 | lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.84 | lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.81 | lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.95 | lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.62 | lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.87 | lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.9 | lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools |
lab-bench-protocolqa-creator-open-response · lab-bench-protocolqa-paper-v3-open-response
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.3 | lab-bench-protocolqa-paper-v3-open-response |
| GPT-4o (LAB-Bench snapshot not reported) | 0.2 | lab-bench-protocolqa-paper-v3-open-response |
lab-bench-seqqa-anthropic-sonnet45-system-card · lab-bench-seqqa-anthropic-10shot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude Sonnet 4 | 0.682 | lab-bench-seqqa-anthropic-10shot-no-tools |
| Claude Opus 4 | 0.723 | lab-bench-seqqa-anthropic-10shot-no-tools |
| Claude Opus 4.1 | 0.785 | lab-bench-seqqa-anthropic-10shot-no-tools |
| Claude Sonnet 4.5 | 0.78 | lab-bench-seqqa-anthropic-10shot-no-tools |
lab-bench-suppqa-creator-mcq · lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.02 | lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.02 | lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.01 | lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.13 | lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.06 | lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.18 | lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.2 | lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools |
lab-bench-suppqa-creator-mcq · lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.75 | lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.59 | lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.32 | lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.47 | lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.39 | lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.35 | lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.27 | lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools |
lab-bench-suppqa-creator-mcq · lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.04 | lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.04 | lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.04 | lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.29 | lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.14 | lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.53 | lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.74 | lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools |
lab-bench-tableqa-creator-mcq · lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.83 | lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.67 | lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.59 | lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.71 | lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.54 | lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.49 | lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools |
lab-bench-tableqa-creator-mcq · lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.9 | lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.74 | lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.71 | lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.75 | lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.58 | lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.51 | lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools |
lab-bench-tableqa-creator-mcq · lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.92 | lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.9 | lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.83 | lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.95 | lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.93 | lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.96 | lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools |
Source locators remain visible; expand an item to inspect the exact Registry fields it supports.
/name/aliases/summary/kind/organizations/release_date/domains/capabilities/modalities/task_counts/total/task_counts/basis/task_counts/subsets/task_counts/subsets/10/count/coverage_notes/access/level/access/tasks/access/artifacts/access/grader/access/license/access/biosafety_notes/versions/0/task_counts/total/versions/0/task_counts/basis/versions/0/task_counts/subsets/scientific_task_classification/entries/1/scientific_task_classification/entries/2/scientific_task_classification/entries/3/scientific_task_classification/entries/4/latest_version/task_counts/total/task_counts/basis/task_counts/subsets/task_counts/subsets/10/count/access/level/access/tasks/access/artifacts/access/grader/access/license/resources/implementations/versions/1/task_counts/total/versions/1/task_counts/basis/versions/1/task_counts/subsets/versions/1/task_counts/subsets/10/count/scientific_task_classification/entries/0/task_counts/subsets/10/countConflicted · high — The official README says 30 narrower subtasks, while the pinned repository contains 31 versioned split files and the paper reports 31 task/result rows./versions/1/task_counts/subsets/10/countConflicted · high — The current snapshot preserves the README claim of 30 alongside the reproducible count of 31 versioned split files.