track · audited · verified 2026-07-21

LAB-Bench CloningScenarios

Human-hard, multi-step multiple-choice scenarios involving plasmids, DNA fragments, enzymes, and molecular-cloning workflows.

Benchmark definition

What is counted

Version
repository-998a8e0
Total
41 (questions across public and private splits)
Task formats
multiple choice
Capabilities
Experiment planningDesignTroubleshootingScientific reasoning
Modalities
TextDNA or RNA sequence

Version history

VersionStatusRelease / as-ofTotalFormal tracks
paper-v3
lab-bench-cloning-scenarios-paper-v3
active2024-07-1741 (questions across public and private splits)None registered
repository-998a8e0
lab-bench-cloning-scenarios-repository-998a8e0
current2025-09-2741 (questions across public and private splits)None registered

Tracks and subsets

IDCountBasisPartition?Notes
Public split
lab-bench-cloning-scenarios-public
33questionsExclusive & exhaustiveReleased in the official repository and Hugging Face dataset.
Private contamination-monitoring split
lab-bench-cloning-scenarios-private
8questionsExclusive & exhaustiveIDs are published; question contents are withheld.
Creator open-response study
lab-bench-cloning-scenarios-creator-open-response
10modified questions sampled for the creator open-response studyNoA non-exhaustive derived subset; wording was modified to remove multiple-choice framing.

Scientific Task Atlas

Scientific task classification

complete for repository-998a8e0. Single-purpose formal LAB-Bench track.

Scientific taskCoverageCountMappingEvidence
Experiment and protocol planningexplicitly-in-scope41 questions
questions across public and private splits
official-track
high confidence
lab-bench-cloning-scenarios-evidence-paper
Molecular-cloning workflow planning.

Evaluation registry

Works and run settings

A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.

Evaluation run

lab-bench-cloning-scenarios-anthropic-sonnet45-system-card

From Claude Sonnet 4.5 System Card

lab-bench-cloning-scenarios-anthropic-10shot-no-toolsvNot reported

Evaluated models / systems: Claude Opus 4, Claude Opus 4.1, Claude Sonnet 4, Claude Sonnet 4.5

Scopetrack
Shots10
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
Statisticsreported point estimate with plotted error bars
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
LAB-Bench scoreabsoluteproportionNot reportedNot reported

Results

ModelMetricValuen
Claude Sonnet 4LAB-Bench score0.485 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported
Claude Opus 4LAB-Bench score0.545 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported
Claude Opus 4.1LAB-Bench score0.758 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported
Claude Sonnet 4.5LAB-Bench score0.667 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported

Evidence

  • section: Section 9.2.4.4, printed pages 132–133 (Four selected tracks, 10-shot prompting, no search/bioinformatics tools, threshold, and unreported scope/repeats/grader.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • figure: Figure 9.2.4.4.A, printed page 133 (Exact point labels. Protocol Sonnet 4/4.5 labels are represented in the Claude for Life Sciences run to avoid duplicate result rows.) — supports /results

Evaluation run

lab-bench-cloning-scenarios-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported)

Scopefull · n=41
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.28 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.54 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.52 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.27 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.41 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3 Opus (claude-3-opus-20240229)Coverage0.65 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.33 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.3 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.28 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.37 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
GPT-4o (LAB-Bench snapshot not reported)Coverage0.77 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.09 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.36 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.31 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.26 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.29 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.9 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
41

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row CloningScenarios (Accuracy, precision, and coverage. Human row excluded.) — supports /results

Evaluation run

lab-bench-cloning-scenarios-creator-mcq-llama-context

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-cloning-scenarios-paper-v3-llama-context-limitedvpaper-v3

Evaluated models / systems: Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=41
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.24 proportion
Only 25 prompts fit; 16 were counted as insufficient-information responses for accuracy and coverage.
41
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.34 proportion
Only 25 prompts fit; 16 were counted as insufficient-information responses for accuracy and coverage.
41
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.72 proportion
Only 25 prompts fit; 16 were counted as insufficient-information responses for accuracy and coverage.
41

Evidence

  • section: Appendix D.2–D.4 (Llama context handling, 25 prompted items, 16 insufficient-information treatments, three-run metrics.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, CloningScenarios row, Meta-Llama-3-70B-Instruct column (Accuracy, precision, and coverage.) — supports /results

Evaluation run

lab-bench-cloning-scenarios-creator-open-response

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-cloning-scenarios-paper-v3-open-responsevpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), GPT-4o (LAB-Bench snapshot not reported)

Scopesubset · n=10
ShotsNot reported
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderexpert-biologist manual grading against the ideal answer, reviewed by a second expert · human review: yes
Statisticssingle reported accuracy per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / sampled modified questionsexpert judgment against ideal multiple-choice answer

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.2 proportion
Creator Table 5 open-response study.
10
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.2 proportion
Creator Table 5 open-response study.
10

Evidence

  • section: Section 2.4 and Appendix B.2 (Subset construction, modified wording, expert grading, and second review.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Table 5 (lab-bench-cloning-scenarios model accuracy and question count.) — supports /results

Comparable result views

LAB-Bench score

lab-bench-cloning-scenarios-anthropic-sonnet45-system-card · lab-bench-cloning-scenarios-anthropic-10shot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude Sonnet 40.485lab-bench-cloning-scenarios-anthropic-10shot-no-tools
Claude Opus 40.545lab-bench-cloning-scenarios-anthropic-10shot-no-tools
Claude Opus 4.10.758lab-bench-cloning-scenarios-anthropic-10shot-no-tools
Claude Sonnet 4.50.667lab-bench-cloning-scenarios-anthropic-10shot-no-tools

Accuracy

lab-bench-cloning-scenarios-creator-mcq · lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.28lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.27lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.1lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.28lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.09lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.26lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools

Precision (selective accuracy)

lab-bench-cloning-scenarios-creator-mcq · lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.54lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.41lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.33lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.37lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.36lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.29lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools

Coverage

lab-bench-cloning-scenarios-creator-mcq · lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.52lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.65lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.3lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.77lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.31lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.9lab-bench-cloning-scenarios-paper-v3-zero-shot-cot-no-tools

Accuracy

lab-bench-cloning-scenarios-creator-mcq-llama-context · lab-bench-cloning-scenarios-paper-v3-llama-context-limited

CSV ↓
Accessible data table
ModelValueComparability group
Meta-Llama-3-70B-Instruct (Anyscale API)0.24lab-bench-cloning-scenarios-paper-v3-llama-context-limited

Precision (selective accuracy)

lab-bench-cloning-scenarios-creator-mcq-llama-context · lab-bench-cloning-scenarios-paper-v3-llama-context-limited

CSV ↓
Accessible data table
ModelValueComparability group
Meta-Llama-3-70B-Instruct (Anyscale API)0.34lab-bench-cloning-scenarios-paper-v3-llama-context-limited

Coverage

lab-bench-cloning-scenarios-creator-mcq-llama-context · lab-bench-cloning-scenarios-paper-v3-llama-context-limited

CSV ↓
Accessible data table
ModelValueComparability group
Meta-Llama-3-70B-Instruct (Anyscale API)0.72lab-bench-cloning-scenarios-paper-v3-llama-context-limited

Accuracy

lab-bench-cloning-scenarios-creator-open-response · lab-bench-cloning-scenarios-paper-v3-open-response

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.2lab-bench-cloning-scenarios-paper-v3-open-response
GPT-4o (LAB-Bench snapshot not reported)0.2lab-bench-cloning-scenarios-paper-v3-open-response

Evidence and change history

Source locators remain visible; expand an item to inspect the exact Registry fields it supports.

LAB-Bench: Measuring Capabilities of Language Models for Biology Research · section: Table 1; Appendix Tables 2–6; task description for CloningScenarios (Task definition, total count, creator protocol, and result row.) · Supports 17 fields

Open source →

  • /name
  • /organizations
  • /release_date
  • /kind
  • /domains
  • /capabilities
  • /modalities
  • /task_counts/total
  • /task_counts/basis
  • /task_counts/subsets
  • /coverage_notes
  • /access/level
  • /access/license
  • /versions/0/task_counts/total
  • /versions/0/task_counts/basis
  • /versions/0/task_counts/subsets
  • /scientific_task_classification/entries/0
lab-bench-cloning-scenarios-repository-resource · repository-path: CloningScenarios/cloningscenarios-v1-splits.json; CloningScenarios/cloningscenarios-v1-public.jsonl where present; task.py at 998a8e0a40cf116c80e1b0e7a805ebb5fb9fa838 (Public/private/total 33/8/41.) · Supports 13 fields

Open source →

  • /latest_version
  • /task_counts/total
  • /task_counts/basis
  • /task_counts/subsets
  • /access/level
  • /access/tasks
  • /access/artifacts
  • /access/grader
  • /access/license
  • /resources
  • /versions/1/task_counts/total
  • /versions/1/task_counts/basis
  • /versions/1/task_counts/subsets

View source-level modification history on GitHub →