suite · audited-with-caveats · verified 2026-07-21

LAB-Bench

A practical biology-research suite of 2,457 multiple-choice questions across eight broad categories and 31 versioned task files, with public and private contamination-monitoring splits.

+10 more
Audited with caveats: 2 field(s) are marked provisional or conflicted. Warnings are shown next to affected values and these claims are excluded from unqualified comparisons.
Category and release audit: LAB-Bench contains 2,457 questions across 8 broad categories. The pinned split manifests reproduce 1,967 public and 490 private questions across 31 task files; the official README instead says 30 narrower subtasks, so that claim remains Conflicted.

Benchmark definition

What is counted

Version
repository-998a8e0
Total
2457 (multiple-choice questions across the complete public and private creator snapshot)
Task formats
multiple choice
Capabilities
KnowledgeEvidence synthesisRetrievalPredictionClassificationDesignData analysisTool useExperiment planningTroubleshootingScientific reasoning
Modalities
TextPaper or documentTableFigureImageDNA or RNA sequenceProtein sequenceDatabaseWeb

Version history

VersionStatusRelease / as-ofTotalFormal tracks
paper-v3
lab-bench-paper-v3
active2024-07-172457 (multiple-choice questions across the complete public and private creator snapshot)lab-bench-litqa2, lab-bench-suppqa, lab-bench-figqa, lab-bench-tableqa, lab-bench-dbqa, lab-bench-protocolqa, lab-bench-seqqa, lab-bench-cloning-scenarios
repository-998a8e0
lab-bench-repository-998a8e0
current2025-09-272457 (multiple-choice questions across the complete public and private creator snapshot)lab-bench-litqa2, lab-bench-suppqa, lab-bench-figqa, lab-bench-tableqa, lab-bench-dbqa, lab-bench-protocolqa, lab-bench-seqqa, lab-bench-cloning-scenarios

Tracks and subsets

IDCountBasisPartition?Notes
LitQA2
lab-bench-litqa2
248questionsExclusive & exhaustive
SuppQA
lab-bench-suppqa
102questionsExclusive & exhaustive
FigQA
lab-bench-figqa
226questionsExclusive & exhaustive
TableQA
lab-bench-tableqa
305questionsExclusive & exhaustive
DbQA
lab-bench-dbqa
650questionsExclusive & exhaustive
ProtocolQA
lab-bench-protocolqa
135questionsExclusive & exhaustive
SeqQA
lab-bench-seqqa
750questionsExclusive & exhaustive
CloningScenarios
lab-bench-cloning-scenarios
41questionsExclusive & exhaustive
Broad categories
lab-bench-broad-categories
8categoriesNoNot a question count.
Versioned formal task files
lab-bench-formal-task-files
31versioned split filesNoThe 31 files are the six single-task categories, 10 DbQA subtasks, 15 SeqQA subtasks, and CloningScenarios.
Narrower subtasks stated in README
lab-bench-readme-narrower-subtasks
30Conflicted · highsubtasks claimed in READMENoConflicts with the 31 versioned split files and 31 creator-result rows.

Registered child tracks

LAB-Bench CloningScenarios

Human-hard, multi-step multiple-choice scenarios involving plasmids, DNA fragments, enzymes, and molecular-cloning workflows.

4 evaluation run(s)

LAB-Bench DbQA

Database-retrieval category spanning 10 genomics, clinical, protein, regulatory, vaccine-response, and viral-PPI tasks.

0 evaluation run(s)

LAB-Bench FigQA

Multiple-choice interpretation and multi-element reasoning over scientific figures shown without captions or paper context.

5 evaluation run(s)

LAB-Bench LitQA2

Literature-retrieval questions whose answers require findings in full research papers rather than titles or abstracts.

1 evaluation run(s)

LAB-Bench ProtocolQA

Troubleshoots intentionally modified published biological protocols by selecting steps that would repair the stated outcome.

4 evaluation run(s)

LAB-Bench SeqQA

Sequence-comprehension and manipulation category spanning 15 formal tasks involving PCR, restriction digestion, ORFs, translation, GC content, and DNA–protein relationships.

1 evaluation run(s)

LAB-Bench SuppQA

Retrieval and interpretation questions answerable from paper supplementary text or PDF tables.

1 evaluation run(s)

LAB-Bench TableQA

Lookup, calculation, and reasoning questions over table images extracted from scientific papers.

1 evaluation run(s)

Scientific Task Atlas

Scientific task classification

partial for repository-998a8e0. Formal child tracks support several precise mappings; the mixed suite has no exhaustive creator scientific-task taxonomy.

Scientific taskCoverageCountMappingEvidence
Scientific database retrievalexplicitly-in-scope650 questions
Questions across the ten formal DbQA child tasks.
official-track
high confidence
lab-bench-evidence-repository
Count is specific to DbQA and is not added to other task claims.
Scientific evidence interpretationexplicitly-in-scopeNot reported
FigQA, LitQA2, SuppQA, and TableQA questions.
official-track
high confidence
lab-bench-evidence-paper
The root record does not publish this cross-track subtotal.
Experiment and protocol planningexplicitly-in-scope135 questions
ProtocolQA questions across public and private splits.
official-track
high confidence
lab-bench-evidence-paper
CloningScenarios is shown separately and is not included in this count.
Protein-protein interaction predictionexplicitly-in-scope50 questions
Viral PPI formal-task questions.
official-track
high confidence
lab-bench-evidence-paper
Viral-human PPI database-retrieval questions.
Protein sequence designnot-in-scope0 questions
Released LAB-Bench questions.
official-taxonomy
high confidence
lab-bench-evidence-paper
Primer and cloning design are not relabeled as protein sequence design.

Scientific coverage notes

DomainCoverageCountInterpretation
Protein sequenceexplicitly-in-scopeNot reportedSeqQA and two ClinVar DbQA tasks require DNA/protein-sequence reasoning, but the sources do not publish a protein-only question count.
Protein-protein bindingexplicitly-in-scope50The Viral PPI formal task has 50 questions about predicted viral–human protein interactions.
Protein designnot-in-scope0Primer and cloning design are in scope; protein sequence design is not a released LAB-Bench task.

Evaluation registry

Works and run settings

A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.

Evaluation run

lab-bench-cloning-scenarios-anthropic-sonnet45-system-card

From Claude Sonnet 4.5 System Card

lab-bench-cloning-scenarios-anthropic-10shot-no-toolsvNot reported

Evaluated models / systems: Claude Opus 4, Claude Opus 4.1, Claude Sonnet 4, Claude Sonnet 4.5

Scopetrack
Shots10
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
Statisticsreported point estimate with plotted error bars
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
LAB-Bench scoreabsoluteproportionNot reportedNot reported

Results

ModelMetricValuen
Claude Sonnet 4LAB-Bench score0.485 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported
Claude Opus 4LAB-Bench score0.545 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported
Claude Opus 4.1LAB-Bench score0.758 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported
Claude Sonnet 4.5LAB-Bench score0.667 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported

Evidence

  • section: Section 9.2.4.4, printed pages 132–133 (Four selected tracks, 10-shot prompting, no search/bioinformatics tools, threshold, and unreported scope/repeats/grader.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • figure: Figure 9.2.4.4.A, printed page 133 (Exact point labels. Protocol Sonnet 4/4.5 labels are represented in the Claude for Life Sciences run to avoid duplicate result rows.) — supports /results

Evaluation run

lab-bench-cloning-scenarios-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported)

Scopefull · n=41
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.28 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.54 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.52 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.27 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.41 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3 Opus (claude-3-opus-20240229)Coverage0.65 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.33 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.3 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.28 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.37 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
GPT-4o (LAB-Bench snapshot not reported)Coverage0.77 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.09 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.36 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.31 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.26 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.29 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.9 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row CloningScenarios (Accuracy, precision, and coverage. Human row excluded.) — supports /results

Evaluation run

lab-bench-cloning-scenarios-creator-mcq-llama-context

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-cloning-scenarios-paper-v3-llama-context-limitedvpaper-v3

Evaluated models / systems: Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=41
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.24 proportion
Only 25 prompts fit; 16 were counted as insufficient-information responses for accuracy and coverage.
41
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.34 proportion
Only 25 prompts fit; 16 were counted as insufficient-information responses for accuracy and coverage.
41
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.72 proportion
Only 25 prompts fit; 16 were counted as insufficient-information responses for accuracy and coverage.
41

Evidence

  • section: Appendix D.2–D.4 (Llama context handling, 25 prompted items, 16 insufficient-information treatments, three-run metrics.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, CloningScenarios row, Meta-Llama-3-70B-Instruct column (Accuracy, precision, and coverage.) — supports /results

Evaluation run

lab-bench-cloning-scenarios-creator-open-response

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-cloning-scenarios-paper-v3-open-responsevpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), GPT-4o (LAB-Bench snapshot not reported)

Scopesubset · n=10
ShotsNot reported
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderexpert-biologist manual grading against the ideal answer, reviewed by a second expert · human review: yes
Statisticssingle reported accuracy per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / sampled modified questionsexpert judgment against ideal multiple-choice answer

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.2 proportion
Creator Table 5 open-response study.
10
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.2 proportion
Creator Table 5 open-response study.
10

Evidence

  • section: Section 2.4 and Appendix B.2 (Subset construction, modified wording, expert grading, and second review.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Table 5 (lab-bench-cloning-scenarios model accuracy and question count.) — supports /results

Evaluation run

lab-bench-figqa-anthropic-sonnet45-system-card

From Claude Sonnet 4.5 System Card

lab-bench-figqa-anthropic-10shot-no-toolsvNot reported

Evaluated models / systems: Claude Opus 4, Claude Opus 4.1, Claude Sonnet 4, Claude Sonnet 4.5

Scopetrack
Shots10
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
Statisticsreported point estimate with plotted error bars
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
LAB-Bench scoreabsoluteproportionNot reportedNot reported

Results

ModelMetricValuen
Claude Sonnet 4LAB-Bench score0.398 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported
Claude Opus 4LAB-Bench score0.508 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported
Claude Opus 4.1LAB-Bench score0.481 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported
Claude Sonnet 4.5LAB-Bench score0.497 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported

Evidence

  • section: Section 9.2.4.4, printed pages 132–133 (Four selected tracks, 10-shot prompting, no search/bioinformatics tools, threshold, and unreported scope/repeats/grader.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • figure: Figure 9.2.4.4.A, printed page 133 (Exact point labels. Protocol Sonnet 4/4.5 labels are represented in the Claude for Life Sciences run to avoid duplicate result rows.) — supports /results
lab-bench-figqa-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported)

Scopefull · n=226
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.46 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.54 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.85 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.24 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.31 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
Claude 3 Opus (claude-3-opus-20240229)Coverage0.78 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.25 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.34 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.73 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.29 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.3 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
GPT-4o (LAB-Bench snapshot not reported)Coverage0.97 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.25 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.28 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.91 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.23 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.24 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.95 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row FigQA (Accuracy, precision, and coverage. Human row excluded.) — supports /results

Evaluation run

lab-bench-figqa-creator-open-response

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-figqa-paper-v3-open-responsevpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), GPT-4o (LAB-Bench snapshot not reported)

Scopesubset · n=10
ShotsNot reported
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderexpert-biologist manual grading against the ideal answer, reviewed by a second expert · human review: yes
Statisticssingle reported accuracy per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / sampled modified questionsexpert judgment against ideal multiple-choice answer

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.3 proportion
Creator Table 5 open-response study.
10
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.3 proportion
Creator Table 5 open-response study.
10

Evidence

  • section: Section 2.4 and Appendix B.2 (Subset construction, modified wording, expert grading, and second review.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Table 5 (lab-bench-figqa model accuracy and question count.) — supports /results

Evaluation run

lab-bench-figqa-crop-tool

From Claude Sonnet 4.6 System Card

lab-bench-figqa-sonnet46-adaptive-max-crop-tool-five-runsvNot reported

Evaluated models / systems: Claude Opus 4.6, Claude Sonnet 4.5, Claude Sonnet 4.6

Scopetrack
ShotsNot reported
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortadaptive thinking at max effort
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolssimple image cropping tool
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats5
GraderNot reported
Statisticsmean over five runs with 95% CI shown
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
FigQA scoreabsolutepercentmean over five runsNot reported

Results

ModelMetricValuen
Claude Sonnet 4.6FigQA score77.1 percent
95% CI is plotted but numeric bounds are not reported.
Not reported
Claude Sonnet 4.5FigQA score59.3 percent
95% CI is plotted but numeric bounds are not reported.
Not reported
Claude Opus 4.6FigQA score78.3 percent
95% CI is plotted but numeric bounds are not reported.
Not reported

Evidence

  • section: Section 2.17.1, printed page 33 (Adaptive thinking, max effort, simple crop tool, five runs, 95% CI, and unreported n/snapshot/grader.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • section: Section 2.17.1 and Figure 2.17.1.A, printed page 33 (All three crop-tool point estimates.) — supports /results

Evaluation run

lab-bench-figqa-no-tools

From Claude Sonnet 4.6 System Card

lab-bench-figqa-sonnet46-adaptive-max-no-tools-five-runsvNot reported

Evaluated models / systems: Claude Opus 4.6, Claude Sonnet 4.5, Claude Sonnet 4.6

Scopetrack
ShotsNot reported
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortadaptive thinking at max effort
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats5
GraderNot reported
Statisticsmean over five runs with 95% CI shown
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
FigQA scoreabsolutepercentmean over five runsNot reported

Results

ModelMetricValuen
Claude Sonnet 4.6FigQA score58.8 percent
95% CI is plotted but numeric bounds are not reported.
Not reported
Claude Sonnet 4.5FigQA score53.4 percent
95% CI is plotted but numeric bounds are not reported.
Not reported
Claude Opus 4.6FigQA score58 percent
95% CI is plotted but numeric bounds are not reported.
Not reported

Evidence

  • section: Section 2.17.1, printed page 33 (Adaptive thinking, max effort, no tools, five runs, 95% CI, and unreported n/snapshot/grader.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • section: Section 2.17.1 and Figure 2.17.1.A, printed page 33 (All three no-tool point estimates.) — supports /results
lab-bench-litqa2-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=248
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.06 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.47 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.12 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.06 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.43 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Claude 3 Opus (claude-3-opus-20240229)Coverage0.14 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.44 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.23 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.27 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.46 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
GPT-4o (LAB-Bench snapshot not reported)Coverage0.58 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.12 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.4 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.31 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.27 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.38 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.7 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.35 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.38 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.92 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row LitQA2 (Accuracy, precision, and coverage. Human row excluded.) — supports /results

Evaluation run

lab-bench-protocolqa-anthropic

From Claude for Life Sciences

lab-bench-protocolqa-anthropic-life-sciences-10shot-no-toolsvNot reported

Evaluated models / systems: Claude Sonnet 4, Claude Sonnet 4.5

Scopetrack
Shots10
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
Statisticsreported point estimate with plotted error bars
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
ProtocolQA scoreabsoluteproportionNot reportedNot reported

Results

ModelMetricValuen
Claude Sonnet 4.5ProtocolQA score0.83 proportion
Official release rounds the system-card point label 0.833 to 0.83.
Not reported
Claude Sonnet 4ProtocolQA score0.74 proportion
Official release rounds the system-card point label 0.741 to 0.74.
Not reported

Evidence

  • section: Section 9.2.4.4 and Figure 9.2.4.4.A, printed pages 132–133 (10-shot, no search/bioinformatics tools, point estimates with error bars, and unreported n/repeats/grader.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • section: Making Claude a better research partner; footnote 1 (Sonnet 4.5 score 0.83, Sonnet 4 score 0.74, and 10-shot prompting.) — supports /results

Evaluation run

lab-bench-protocolqa-anthropic-sonnet45-system-card

From Claude Sonnet 4.5 System Card

lab-bench-protocolqa-anthropic-10shot-no-toolsvNot reported

Evaluated models / systems: Claude Opus 4, Claude Opus 4.1

Scopetrack
Shots10
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
Statisticsreported point estimate with plotted error bars
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
LAB-Bench scoreabsoluteproportionNot reportedNot reported

Results

ModelMetricValuen
Claude Opus 4LAB-Bench score0.796 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported
Claude Opus 4.1LAB-Bench score0.833 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported

Evidence

  • section: Section 9.2.4.4, printed pages 132–133 (Four selected tracks, 10-shot prompting, no search/bioinformatics tools, threshold, and unreported scope/repeats/grader.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • figure: Figure 9.2.4.4.A, printed page 133 (Exact point labels. Protocol Sonnet 4/4.5 labels are represented in the Claude for Life Sciences run to avoid duplicate result rows.) — supports /results

Evaluation run

lab-bench-protocolqa-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-protocolqa-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=135
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.48 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.66 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.73 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.52 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.62 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Claude 3 Opus (claude-3-opus-20240229)Coverage0.84 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.47 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.58 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.81 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.53 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.56 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
GPT-4o (LAB-Bench snapshot not reported)Coverage0.95 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.37 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.59 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.62 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.45 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.51 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.87 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.44 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.49 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.9 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row ProtocolQA (Accuracy, precision, and coverage. Human row excluded.) — supports /results

Evaluation run

lab-bench-protocolqa-creator-open-response

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-protocolqa-paper-v3-open-responsevpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), GPT-4o (LAB-Bench snapshot not reported)

Scopesubset · n=20
ShotsNot reported
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderexpert-biologist manual grading against the ideal answer, reviewed by a second expert · human review: yes
Statisticssingle reported accuracy per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / sampled modified questionsexpert judgment against ideal multiple-choice answer

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.3 proportion
Creator Table 5 open-response study.
20
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.2 proportion
Creator Table 5 open-response study.
20

Evidence

  • section: Section 2.4 and Appendix B.2 (Subset construction, modified wording, expert grading, and second review.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Table 5 (lab-bench-protocolqa model accuracy and question count.) — supports /results

Evaluation run

lab-bench-seqqa-anthropic-sonnet45-system-card

From Claude Sonnet 4.5 System Card

lab-bench-seqqa-anthropic-10shot-no-toolsvNot reported

Evaluated models / systems: Claude Opus 4, Claude Opus 4.1, Claude Sonnet 4, Claude Sonnet 4.5

Scopetrack
Shots10
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
Statisticsreported point estimate with plotted error bars
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
LAB-Bench scoreabsoluteproportionNot reportedNot reported

Results

ModelMetricValuen
Claude Sonnet 4LAB-Bench score0.682 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported
Claude Opus 4LAB-Bench score0.723 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported
Claude Opus 4.1LAB-Bench score0.785 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported
Claude Sonnet 4.5LAB-Bench score0.78 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported

Evidence

  • section: Section 9.2.4.4, printed pages 132–133 (Four selected tracks, 10-shot prompting, no search/bioinformatics tools, threshold, and unreported scope/repeats/grader.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • figure: Figure 9.2.4.4.A, printed page 133 (Exact point labels. Protocol Sonnet 4/4.5 labels are represented in the Claude for Life Sciences run to avoid duplicate result rows.) — supports /results
lab-bench-suppqa-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=102
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.02 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.75 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.04 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.02 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.59 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Claude 3 Opus (claude-3-opus-20240229)Coverage0.04 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.01 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.32 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.04 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.13 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.47 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
GPT-4o (LAB-Bench snapshot not reported)Coverage0.29 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.06 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.39 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.14 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.18 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.35 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.53 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.2 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.27 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.74 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row SuppQA (Accuracy, precision, and coverage. Human row excluded.) — supports /results
lab-bench-tableqa-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported)

Scopefull · n=305
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.83 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.9 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.92 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.67 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.74 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
Claude 3 Opus (claude-3-opus-20240229)Coverage0.9 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.59 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.71 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.83 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.71 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.75 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
GPT-4o (LAB-Bench snapshot not reported)Coverage0.95 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.54 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.58 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.93 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.49 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.51 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.96 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row TableQA (Accuracy, precision, and coverage. Human row excluded.) — supports /results

Comparable result views

LAB-Bench score

lab-bench-cloning-scenarios-anthropic-sonnet45-system-card · lab-bench-cloning-scenarios-anthropic-10shot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude Sonnet 40.485lab-bench-cloning-scenarios-anthropic-10shot-no-tools
Claude Opus 40.545lab-bench-cloning-scenarios-anthropic-10shot-no-tools
Claude Opus 4.10.758lab-bench-cloning-scenarios-anthropic-10shot-no-tools
Claude Sonnet 4.50.667lab-bench-cloning-scenarios-anthropic-10shot-no-tools

Accuracy

lab-bench-cloning-scenarios-creator-mcq · lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.28lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.27lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.1lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.28lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.09lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.26lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools

Precision (selective accuracy)

lab-bench-cloning-scenarios-creator-mcq · lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.54lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.41lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.33lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.37lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.36lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.29lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools

Coverage

lab-bench-cloning-scenarios-creator-mcq · lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.52lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.65lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.3lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.77lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.31lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.9lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools

Accuracy

lab-bench-cloning-scenarios-creator-mcq-llama-context · lab-bench-cloning-scenarios-paper-v3-llama-context-limited

CSV ↓
Accessible data table
ModelValueComparability group
Meta-Llama-3-70B-Instruct (Anyscale API)0.24lab-bench-cloning-scenarios-paper-v3-llama-context-limited

Precision (selective accuracy)

lab-bench-cloning-scenarios-creator-mcq-llama-context · lab-bench-cloning-scenarios-paper-v3-llama-context-limited

CSV ↓
Accessible data table
ModelValueComparability group
Meta-Llama-3-70B-Instruct (Anyscale API)0.34lab-bench-cloning-scenarios-paper-v3-llama-context-limited

Coverage

lab-bench-cloning-scenarios-creator-mcq-llama-context · lab-bench-cloning-scenarios-paper-v3-llama-context-limited

CSV ↓
Accessible data table
ModelValueComparability group
Meta-Llama-3-70B-Instruct (Anyscale API)0.72lab-bench-cloning-scenarios-paper-v3-llama-context-limited

Accuracy

lab-bench-cloning-scenarios-creator-open-response · lab-bench-cloning-scenarios-paper-v3-open-response

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.2lab-bench-cloning-scenarios-paper-v3-open-response
GPT-4o (LAB-Bench snapshot not reported)0.2lab-bench-cloning-scenarios-paper-v3-open-response

LAB-Bench score

lab-bench-figqa-anthropic-sonnet45-system-card · lab-bench-figqa-anthropic-10shot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude Sonnet 40.398lab-bench-figqa-anthropic-10shot-no-tools
Claude Opus 40.508lab-bench-figqa-anthropic-10shot-no-tools
Claude Opus 4.10.481lab-bench-figqa-anthropic-10shot-no-tools
Claude Sonnet 4.50.497lab-bench-figqa-anthropic-10shot-no-tools

Accuracy

lab-bench-figqa-creator-mcq · lab-bench-figqa-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.46lab-bench-figqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.24lab-bench-figqa-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.25lab-bench-figqa-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.29lab-bench-figqa-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.25lab-bench-figqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.23lab-bench-figqa-paper-v3-zero-shot-cot-no-tools

Precision (selective accuracy)

lab-bench-figqa-creator-mcq · lab-bench-figqa-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.54lab-bench-figqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.31lab-bench-figqa-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.34lab-bench-figqa-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.3lab-bench-figqa-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.28lab-bench-figqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.24lab-bench-figqa-paper-v3-zero-shot-cot-no-tools

Coverage

lab-bench-figqa-creator-mcq · lab-bench-figqa-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.85lab-bench-figqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.78lab-bench-figqa-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.73lab-bench-figqa-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.97lab-bench-figqa-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.91lab-bench-figqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.95lab-bench-figqa-paper-v3-zero-shot-cot-no-tools

Accuracy

lab-bench-figqa-creator-open-response · lab-bench-figqa-paper-v3-open-response

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.3lab-bench-figqa-paper-v3-open-response
GPT-4o (LAB-Bench snapshot not reported)0.3lab-bench-figqa-paper-v3-open-response

FigQA score

lab-bench-figqa-crop-tool · lab-bench-figqa-sonnet46-adaptive-max-crop-tool-five-runs

CSV ↓
Accessible data table
ModelValueComparability group
Claude Sonnet 4.677.1lab-bench-figqa-sonnet46-adaptive-max-crop-tool-five-runs
Claude Sonnet 4.559.3lab-bench-figqa-sonnet46-adaptive-max-crop-tool-five-runs
Claude Opus 4.678.3lab-bench-figqa-sonnet46-adaptive-max-crop-tool-five-runs

FigQA score

lab-bench-figqa-no-tools · lab-bench-figqa-sonnet46-adaptive-max-no-tools-five-runs

CSV ↓
Accessible data table
ModelValueComparability group
Claude Sonnet 4.658.8lab-bench-figqa-sonnet46-adaptive-max-no-tools-five-runs
Claude Sonnet 4.553.4lab-bench-figqa-sonnet46-adaptive-max-no-tools-five-runs
Claude Opus 4.658lab-bench-figqa-sonnet46-adaptive-max-no-tools-five-runs

Accuracy

lab-bench-litqa2-creator-mcq · lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.06lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.06lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.1lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.27lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.12lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.27lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.35lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools

Precision (selective accuracy)

lab-bench-litqa2-creator-mcq · lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.47lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.43lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.44lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.46lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.4lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.38lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.38lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools

Coverage

lab-bench-litqa2-creator-mcq · lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.12lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.14lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.23lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.58lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.31lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.7lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.92lab-bench-litqa2-paper-v3-zero-shot-cot-no-tools

ProtocolQA score

lab-bench-protocolqa-anthropic · lab-bench-protocolqa-anthropic-life-sciences-10shot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude Sonnet 4.50.83lab-bench-protocolqa-anthropic-life-sciences-10shot-no-tools
Claude Sonnet 40.74lab-bench-protocolqa-anthropic-life-sciences-10shot-no-tools

LAB-Bench score

lab-bench-protocolqa-anthropic-sonnet45-system-card · lab-bench-protocolqa-anthropic-10shot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude Opus 40.796lab-bench-protocolqa-anthropic-10shot-no-tools
Claude Opus 4.10.833lab-bench-protocolqa-anthropic-10shot-no-tools

Accuracy

lab-bench-protocolqa-creator-mcq · lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.48lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.52lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.47lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.53lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.37lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.45lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.44lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools

Precision (selective accuracy)

lab-bench-protocolqa-creator-mcq · lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.66lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.62lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.58lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.56lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.59lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.51lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.49lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools

Coverage

lab-bench-protocolqa-creator-mcq · lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.73lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.84lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.81lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.95lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.62lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.87lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.9lab-bench-protocolqa-paper-v3-zero-shot-cot-no-tools

Accuracy

lab-bench-protocolqa-creator-open-response · lab-bench-protocolqa-paper-v3-open-response

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.3lab-bench-protocolqa-paper-v3-open-response
GPT-4o (LAB-Bench snapshot not reported)0.2lab-bench-protocolqa-paper-v3-open-response

LAB-Bench score

lab-bench-seqqa-anthropic-sonnet45-system-card · lab-bench-seqqa-anthropic-10shot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude Sonnet 40.682lab-bench-seqqa-anthropic-10shot-no-tools
Claude Opus 40.723lab-bench-seqqa-anthropic-10shot-no-tools
Claude Opus 4.10.785lab-bench-seqqa-anthropic-10shot-no-tools
Claude Sonnet 4.50.78lab-bench-seqqa-anthropic-10shot-no-tools

Accuracy

lab-bench-suppqa-creator-mcq · lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.02lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.02lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.01lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.13lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.06lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.18lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.2lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools

Precision (selective accuracy)

lab-bench-suppqa-creator-mcq · lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.75lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.59lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.32lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.47lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.39lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.35lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.27lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools

Coverage

lab-bench-suppqa-creator-mcq · lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.04lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.04lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.04lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.29lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.14lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.53lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.74lab-bench-suppqa-paper-v3-zero-shot-cot-no-tools

Accuracy

lab-bench-tableqa-creator-mcq · lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.83lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.67lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.59lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.71lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.54lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.49lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools

Precision (selective accuracy)

lab-bench-tableqa-creator-mcq · lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.9lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.74lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.71lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.75lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.58lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.51lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools

Coverage

lab-bench-tableqa-creator-mcq · lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.92lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.9lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.83lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.95lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.93lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.96lab-bench-tableqa-paper-v3-zero-shot-cot-no-tools

Evidence and change history

Source locators remain visible; expand an item to inspect the exact Registry fields it supports.

LAB-Bench: Measuring Capabilities of Language Models for Biology Research · section: Abstract; Sections 1–2; Table 1; Appendix B–E; Tables 2–6 (Definition, full count, category/subtask rows, public/private policy, task descriptions, creator protocol, metrics, models, and results.) · Supports 27 fields

Open source →

  • /name
  • /aliases
  • /summary
  • /kind
  • /organizations
  • /release_date
  • /domains
  • /capabilities
  • /modalities
  • /task_counts/total
  • /task_counts/basis
  • /task_counts/subsets
  • /task_counts/subsets/10/count
  • /coverage_notes
  • /access/level
  • /access/tasks
  • /access/artifacts
  • /access/grader
  • /access/license
  • /access/biosafety_notes
  • /versions/0/task_counts/total
  • /versions/0/task_counts/basis
  • /versions/0/task_counts/subsets
  • /scientific_task_classification/entries/1
  • /scientific_task_classification/entries/2
  • /scientific_task_classification/entries/3
  • /scientific_task_classification/entries/4
lab-bench-repository-resource · repository-path: README.md; CHANGELOG.md; LICENSE; 31 *-splits.json files; labbench/evaluator.py at 998a8e0a40cf116c80e1b0e7a805ebb5fb9fa838 (31 files sum to public/private/total 1,967/490/2,457; README states 8 categories and 30 narrower subtasks.) · Supports 17 fields

Open source →

  • /latest_version
  • /task_counts/total
  • /task_counts/basis
  • /task_counts/subsets
  • /task_counts/subsets/10/count
  • /access/level
  • /access/tasks
  • /access/artifacts
  • /access/grader
  • /access/license
  • /resources
  • /implementations
  • /versions/1/task_counts/total
  • /versions/1/task_counts/basis
  • /versions/1/task_counts/subsets
  • /versions/1/task_counts/subsets/10/count
  • /scientific_task_classification/entries/0

Unresolved field claims

  • /task_counts/subsets/10/countConflicted · high — The official README says 30 narrower subtasks, while the pinned repository contains 31 versioned split files and the paper reports 31 task/result rows.
    Evidence: lab-bench-evidence-paper, lab-bench-evidence-repository
  • /versions/1/task_counts/subsets/10/countConflicted · high — The current snapshot preserves the README claim of 30 alongside the reproducible count of 31 versioned split files.
    Evidence: lab-bench-evidence-repository

View source-level modification history on GitHub →