paper · benchmark creator

LAB-Bench: Measuring Capabilities of Language Models for Biology Research

FutureHouse · 2024-07-14

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

No BenchmarkUse relation is normalized for this legacy work yet. Existing EvaluationRuns remain available below.

Normalized evaluation runs

LAB-Bench CloningScenarios3 runs

Open benchmark record →

Evaluation run

lab-bench-cloning-scenarios-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported)

Scopefull · n=41
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.28 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.54 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.52 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.27 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.41 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3 Opus (claude-3-opus-20240229)Coverage0.65 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.33 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.3 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.28 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.37 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
GPT-4o (LAB-Bench snapshot not reported)Coverage0.77 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.09 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.36 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.31 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.26 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.29 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.9 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row CloningScenarios (Accuracy, precision, and coverage. Human row excluded.) — supports /results

Evaluation run

lab-bench-cloning-scenarios-creator-mcq-llama-context

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-cloning-scenarios-paper-v3-llama-context-limitedvpaper-v3

Evaluated models / systems: Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=41
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.24 proportion
Only 25 prompts fit; 16 were counted as insufficient-information responses for accuracy and coverage.
41
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.34 proportion
Only 25 prompts fit; 16 were counted as insufficient-information responses for accuracy and coverage.
41
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.72 proportion
Only 25 prompts fit; 16 were counted as insufficient-information responses for accuracy and coverage.
41

Evidence

  • section: Appendix D.2–D.4 (Llama context handling, 25 prompted items, 16 insufficient-information treatments, three-run metrics.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, CloningScenarios row, Meta-Llama-3-70B-Instruct column (Accuracy, precision, and coverage.) — supports /results

Evaluation run

lab-bench-cloning-scenarios-creator-open-response

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-cloning-scenarios-paper-v3-open-responsevpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), GPT-4o (LAB-Bench snapshot not reported)

Scopesubset · n=10
ShotsNot reported
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderexpert-biologist manual grading against the ideal answer, reviewed by a second expert · human review: yes
Statisticssingle reported accuracy per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / sampled modified questionsexpert judgment against ideal multiple-choice answer

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.2 proportion
Creator Table 5 open-response study.
10
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.2 proportion
Creator Table 5 open-response study.
10

Evidence

  • section: Section 2.4 and Appendix B.2 (Subset construction, modified wording, expert grading, and second review.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Table 5 (lab-bench-cloning-scenarios model accuracy and question count.) — supports /results
LAB-Bench DbQA — Disease gene associations1 run

Open benchmark record →

Evaluation run

lab-bench-dbqa-dga-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0.01 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.16 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.17 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.95 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.05 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.17 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.32 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.15 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.17 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.9 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.25 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.25 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.99 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row DbQA_dga_task-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench DbQA — Gene location1 run

Open benchmark record →

Evaluation run

lab-bench-dbqa-gene-location-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0.01 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.03 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.22 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.09 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.31 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.38 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.81 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.12 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.22 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.55 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.18 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.22 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.81 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.2 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.21 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.97 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row DbQA_gene_location_task-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench DbQA — miRNA targets1 run

Open benchmark record →

Evaluation run

lab-bench-dbqa-mirna-targets-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.3 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.35 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.01 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.03 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.15 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.15 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.26 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.6 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.29 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.29 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.99 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row DbQA_mirna_targets_task-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench DbQA — Mouse tumor gene sets1 run

Open benchmark record →

Evaluation run

lab-bench-dbqa-mouse-tumor-gene-sets-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=100
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.55 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.73 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.76 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.36 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.71 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Opus (claude-3-opus-20240229)Coverage0.51 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.76 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.14 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.57 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.59 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4o (LAB-Bench snapshot not reported)Coverage0.96 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.16 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.58 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.28 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.48 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.55 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.87 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.53 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.58 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.93 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row DbQA_mouse_tumor_gene_sets-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench DbQA — Oncogenic signatures1 run

Open benchmark record →

Evaluation run

lab-bench-dbqa-oncogenic-signatures-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.14 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.59 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.24 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.12 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.57 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0.2 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.04 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.75 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.05 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.25 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.32 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.79 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.06 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.61 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.24 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.3 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.8 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.2 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.47 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.42 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row DbQA_oncogenic_signatures_task-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench DbQA — GTRD transcription-factor binding sites1 run

Open benchmark record →

Evaluation run

lab-bench-dbqa-tfbs-gtrd-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.01 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.01 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.33 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.01 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.11 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.11 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.31 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.35 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.19 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.19 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row DbQA_tfbs_GTRD_task-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench DbQA — Protein variant from sequence1 run

Open benchmark record →

Evaluation run

lab-bench-dbqa-variant-from-sequence-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=100
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.02 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.75 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.03 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Opus (claude-3-opus-20240229)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.11 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Opus (claude-3-opus-20240229)Coverage0.03 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.17 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.01 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.42 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.44 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4o (LAB-Bench snapshot not reported)Coverage0.95 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.01 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.22 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.36 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.62 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.3 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.35 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.85 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row DbQA_variant_from_sequence_task-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench DbQA — Protein variant with multiple sequences1 run

Open benchmark record →

Evaluation run

lab-bench-dbqa-variant-multi-sequence-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=100
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.09 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.16 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.56 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.05 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.18 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Opus (claude-3-opus-20240229)Coverage0.26 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.33 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.04 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.06 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.12 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4o (LAB-Bench snapshot not reported)Coverage0.49 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.02 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.14 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.16 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.84 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.07 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.12 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.63 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row DbQA_variant_multi_sequence_task-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench DbQA — Vaccine response gene sets1 run

Open benchmark record →

Evaluation run

lab-bench-dbqa-vax-response-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.21 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.57 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.36 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.14 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.65 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0.21 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.13 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.61 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.21 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.33 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.45 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.73 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.06 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.65 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.09 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.32 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.38 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.83 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.15 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.54 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.29 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row DbQA_vax_response_task-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench DbQA — Viral protein–protein interactions1 run

Open benchmark record →

Evaluation run

lab-bench-dbqa-viral-ppi-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.26 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.69 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.38 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.02 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.28 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.06 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.27 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.49 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.55 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.4 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.4 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row DbQA_viral_ppi_task-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench FigQA2 runs

Open benchmark record →

lab-bench-figqa-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported)

Scopefull · n=226
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.46 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.54 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.85 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.24 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.31 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
Claude 3 Opus (claude-3-opus-20240229)Coverage0.78 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.25 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.34 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.73 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.29 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.3 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
GPT-4o (LAB-Bench snapshot not reported)Coverage0.97 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.25 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.28 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.91 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.23 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.24 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.95 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
226

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row FigQA (Accuracy, precision, and coverage. Human row excluded.) — supports /results

Evaluation run

lab-bench-figqa-creator-open-response

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-figqa-paper-v3-open-responsevpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), GPT-4o (LAB-Bench snapshot not reported)

Scopesubset · n=10
ShotsNot reported
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderexpert-biologist manual grading against the ideal answer, reviewed by a second expert · human review: yes
Statisticssingle reported accuracy per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / sampled modified questionsexpert judgment against ideal multiple-choice answer

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.3 proportion
Creator Table 5 open-response study.
10
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.3 proportion
Creator Table 5 open-response study.
10

Evidence

  • section: Section 2.4 and Appendix B.2 (Subset construction, modified wording, expert grading, and second review.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Table 5 (lab-bench-figqa model accuracy and question count.) — supports /results
LAB-Bench LitQA21 run

Open benchmark record →

lab-bench-litqa2-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=248
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.06 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.47 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.12 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.06 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.43 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Claude 3 Opus (claude-3-opus-20240229)Coverage0.14 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.44 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.23 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.27 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.46 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
GPT-4o (LAB-Bench snapshot not reported)Coverage0.58 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.12 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.4 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.31 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.27 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.38 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.7 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.35 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.38 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.92 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
248

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row LitQA2 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench ProtocolQA2 runs

Open benchmark record →

Evaluation run

lab-bench-protocolqa-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-protocolqa-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=135
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.48 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.66 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.73 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.52 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.62 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Claude 3 Opus (claude-3-opus-20240229)Coverage0.84 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.47 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.58 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.81 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.53 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.56 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
GPT-4o (LAB-Bench snapshot not reported)Coverage0.95 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.37 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.59 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.62 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.45 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.51 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.87 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.44 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.49 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.9 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
135

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row ProtocolQA (Accuracy, precision, and coverage. Human row excluded.) — supports /results

Evaluation run

lab-bench-protocolqa-creator-open-response

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-protocolqa-paper-v3-open-responsevpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), GPT-4o (LAB-Bench snapshot not reported)

Scopesubset · n=20
ShotsNot reported
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderexpert-biologist manual grading against the ideal answer, reviewed by a second expert · human review: yes
Statisticssingle reported accuracy per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / sampled modified questionsexpert judgment against ideal multiple-choice answer

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.3 proportion
Creator Table 5 open-response study.
20
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.2 proportion
Creator Table 5 open-response study.
20

Evidence

  • section: Section 2.4 and Appendix B.2 (Subset construction, modified wording, expert grading, and second review.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Table 5 (lab-bench-protocolqa model accuracy and question count.) — supports /results
LAB-Bench SeqQA — ORF amino-acid position1 run

Open benchmark record →

Evaluation run

lab-bench-seqqa-orf-seq-aaid-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-seqqa-orf-seq-aaid-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.07 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.13 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.57 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.03 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.13 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0.2 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.02 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.09 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.23 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.16 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.23 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.69 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.07 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.15 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.47 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.17 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.19 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.91 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.19 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.19 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.97 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row SeqQA_ORF-seq-AAid-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench SeqQA — ORF amino-acid sequence1 run

Open benchmark record →

Evaluation run

lab-bench-seqqa-orf-seq-aaseq-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-seqqa-orf-seq-aaseq-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.87 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.89 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.99 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.67 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.7 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0.95 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.41 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.67 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.61 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.5 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.5 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.99 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.39 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.5 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.77 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.57 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.57 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.49 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.49 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row SeqQA_ORF-seq-AAseq-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench SeqQA — ORF count above length1 run

Open benchmark record →

Evaluation run

lab-bench-seqqa-orf-seq-numlen-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-seqqa-orf-seq-numlen-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.01 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.07 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.28 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0.26 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.05 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.18 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.31 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.11 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.14 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.81 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.07 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.27 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.29 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.26 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.26 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.23 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.29 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.81 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row SeqQA_ORF-seq-numlen-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench SeqQA — Translation efficiency1 run

Open benchmark record →

Evaluation run

lab-bench-seqqa-orf-transeff-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-seqqa-orf-transeff-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.87 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.92 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.95 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.64 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.71 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0.9 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.47 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.67 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.71 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.48 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.66 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.73 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.29 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.49 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.59 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.47 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.52 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.92 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.45 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.46 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.98 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row SeqQA_ORF-transeff-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench SeqQA — Gene-to-restriction primers1 run

Open benchmark record →

Evaluation run

lab-bench-seqqa-pcr-gene-enzprimers-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-seqqa-pcr-gene-enzprimers-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.57 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.78 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.72 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.75 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.87 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0.86 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.47 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.8 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.59 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.6 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.64 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.93 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.49 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.6 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.81 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.61 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.61 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.58 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.58 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.99 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row SeqQA_PCR-gene-enzprimers-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench SeqQA — Gene-to-Gibson primers (HindIII)1 run

Open benchmark record →

Evaluation run

lab-bench-seqqa-pcr-gene-gibshindprimers-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-seqqa-pcr-gene-gibshindprimers-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.61 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.71 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.85 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.48 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.75 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0.65 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.43 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.71 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.61 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.64 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.75 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.85 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.13 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.67 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.19 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.57 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.57 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.45 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.45 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row SeqQA_PCR-gene-gibshindprimers-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench SeqQA — Gene-to-Gibson primers (SmaI)1 run

Open benchmark record →

Evaluation run

lab-bench-seqqa-pcr-gene-gibssmaprimers-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-seqqa-pcr-gene-gibssmaprimers-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.65 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.7 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.93 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.34 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.4 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0.86 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.16 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.54 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.3 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.45 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.47 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.95 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.29 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.72 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.4 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.51 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.51 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.99 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.58 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.58 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.99 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row SeqQA_PCR-gene-gibssmaprimers-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench SeqQA — Primers-to-restriction enzymes1 run

Open benchmark record →

Evaluation run

lab-bench-seqqa-pcr-geneprimers-enz-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-seqqa-pcr-geneprimers-enz-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.73 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.91 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.81 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.76 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.84 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0.91 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.66 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.69 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.95 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.74 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.77 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.97 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.66 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.75 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.89 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.83 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.83 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.37 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.39 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.95 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row SeqQA_PCR-geneprimers-enz-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench SeqQA — Amplicon length to primers1 run

Open benchmark record →

Evaluation run

lab-bench-seqqa-pcr-len-primers-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-seqqa-pcr-len-primers-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.29 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.29 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.99 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.37 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.45 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0.81 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.02 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.07 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.31 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.12 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.14 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.84 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.13 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.78 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.13 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.14 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.95 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.06 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.06 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row SeqQA_PCR-len-primers-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench SeqQA — Primers to amplicon length1 run

Open benchmark record →

Evaluation run

lab-bench-seqqa-pcr-primers-len-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-seqqa-pcr-primers-len-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.35 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.37 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.93 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.29 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.38 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0.78 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.15 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.19 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.77 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.13 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.15 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.89 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.19 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.26 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.71 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.26 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.26 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.21 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.22 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.97 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row SeqQA_PCR-primers-len-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench SeqQA — Sequence-to-restriction primers1 run

Open benchmark record →

Evaluation run

lab-bench-seqqa-pcr-seq-enzprimers-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-seqqa-pcr-seq-enzprimers-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.94 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.95 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.99 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.8 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.83 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0.96 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.54 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.73 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.75 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.83 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.86 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.96 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.67 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.78 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.87 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.86 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.86 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.73 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.73 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row SeqQA_PCR-seq-enzprimers-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench SeqQA — Amplicon sequence to primers1 run

Open benchmark record →

Evaluation run

lab-bench-seqqa-pcr-seq-primers-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-seqqa-pcr-seq-primers-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.97 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.98 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.99 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.83 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.91 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0.91 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.48 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.64 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.75 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.73 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.76 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.96 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.52 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.56 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.93 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.81 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.82 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.99 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.47 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.47 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row SeqQA_PCR-seq-primers-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench SeqQA — GC percentage1 run

Open benchmark record →

Evaluation run

lab-bench-seqqa-prop-seq-gcpercent-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-seqqa-prop-seq-gcpercent-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.59 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.59 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.99 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.45 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.5 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0.9 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.39 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.41 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.93 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.29 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.32 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.9 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.18 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.23 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.78 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.39 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.4 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.99 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.35 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.36 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.99 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row SeqQA_Prop-seq-gcpercent-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench SeqQA — Restriction-fragment lengths1 run

Open benchmark record →

Evaluation run

lab-bench-seqqa-re-seq-lenfrags-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-seqqa-re-seq-lenfrags-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.12 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.33 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.36 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.13 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.26 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0.51 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.07 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.17 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.43 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.17 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.25 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.67 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.07 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.27 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.29 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.21 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.22 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.97 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.19 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.19 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row SeqQA_RE-seq-lenfrags-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench SeqQA — Restriction-fragment count1 run

Open benchmark record →

Evaluation run

lab-bench-seqqa-re-seq-numfrags-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-seqqa-re-seq-numfrags-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.28 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.35 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.13 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.18 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0.76 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.17 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.22 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.75 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.09 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.11 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.81 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.12 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.22 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.54 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.12 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.12 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.97 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.15 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.16 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.96 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row SeqQA_RE-seq-numfrags-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench SuppQA1 run

Open benchmark record →

lab-bench-suppqa-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=102
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.02 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.75 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.04 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.02 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.59 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Claude 3 Opus (claude-3-opus-20240229)Coverage0.04 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.01 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.32 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.04 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.13 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.47 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
GPT-4o (LAB-Bench snapshot not reported)Coverage0.29 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.06 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.39 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.14 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.18 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.35 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.53 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.2 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.27 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.74 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
102

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row SuppQA (Accuracy, precision, and coverage. Human row excluded.) — supports /results
LAB-Bench TableQA1 run

Open benchmark record →

lab-bench-tableqa-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported)

Scopefull · n=305
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.83 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.9 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.92 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.67 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.74 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
Claude 3 Opus (claude-3-opus-20240229)Coverage0.9 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.59 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.71 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.83 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.71 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.75 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
GPT-4o (LAB-Bench snapshot not reported)Coverage0.95 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.54 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.58 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.93 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.49 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.51 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.96 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
305

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row TableQA (Accuracy, precision, and coverage. Human row excluded.) — supports /results