system-card · official model provider

Claude Sonnet 4.5 System Card

Anthropic · 2025-09-29

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

No BenchmarkUse relation is normalized for this legacy work yet. Existing EvaluationRuns remain available below.

Normalized evaluation runs

LAB-Bench CloningScenarios1 run

Open benchmark record →

Evaluation run

lab-bench-cloning-scenarios-anthropic-sonnet45-system-card

From Claude Sonnet 4.5 System Card

lab-bench-cloning-scenarios-anthropic-10shot-no-toolsvNot reported

Evaluated models / systems: Claude Opus 4, Claude Opus 4.1, Claude Sonnet 4, Claude Sonnet 4.5

Scopetrack
Shots10
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
Statisticsreported point estimate with plotted error bars
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
LAB-Bench scoreabsoluteproportionNot reportedNot reported

Results

ModelMetricValuen
Claude Sonnet 4LAB-Bench score0.485 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported
Claude Opus 4LAB-Bench score0.545 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported
Claude Opus 4.1LAB-Bench score0.758 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported
Claude Sonnet 4.5LAB-Bench score0.667 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported

Evidence

  • section: Section 9.2.4.4, printed pages 132–133 (Four selected tracks, 10-shot prompting, no search/bioinformatics tools, threshold, and unreported scope/repeats/grader.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • figure: Figure 9.2.4.4.A, printed page 133 (Exact point labels. Protocol Sonnet 4/4.5 labels are represented in the Claude for Life Sciences run to avoid duplicate result rows.) — supports /results
LAB-Bench FigQA1 run

Open benchmark record →

Evaluation run

lab-bench-figqa-anthropic-sonnet45-system-card

From Claude Sonnet 4.5 System Card

lab-bench-figqa-anthropic-10shot-no-toolsvNot reported

Evaluated models / systems: Claude Opus 4, Claude Opus 4.1, Claude Sonnet 4, Claude Sonnet 4.5

Scopetrack
Shots10
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
Statisticsreported point estimate with plotted error bars
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
LAB-Bench scoreabsoluteproportionNot reportedNot reported

Results

ModelMetricValuen
Claude Sonnet 4LAB-Bench score0.398 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported
Claude Opus 4LAB-Bench score0.508 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported
Claude Opus 4.1LAB-Bench score0.481 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported
Claude Sonnet 4.5LAB-Bench score0.497 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported

Evidence

  • section: Section 9.2.4.4, printed pages 132–133 (Four selected tracks, 10-shot prompting, no search/bioinformatics tools, threshold, and unreported scope/repeats/grader.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • figure: Figure 9.2.4.4.A, printed page 133 (Exact point labels. Protocol Sonnet 4/4.5 labels are represented in the Claude for Life Sciences run to avoid duplicate result rows.) — supports /results
LAB-Bench ProtocolQA1 run

Open benchmark record →

Evaluation run

lab-bench-protocolqa-anthropic-sonnet45-system-card

From Claude Sonnet 4.5 System Card

lab-bench-protocolqa-anthropic-10shot-no-toolsvNot reported

Evaluated models / systems: Claude Opus 4, Claude Opus 4.1

Scopetrack
Shots10
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
Statisticsreported point estimate with plotted error bars
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
LAB-Bench scoreabsoluteproportionNot reportedNot reported

Results

ModelMetricValuen
Claude Opus 4LAB-Bench score0.796 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported
Claude Opus 4.1LAB-Bench score0.833 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported

Evidence

  • section: Section 9.2.4.4, printed pages 132–133 (Four selected tracks, 10-shot prompting, no search/bioinformatics tools, threshold, and unreported scope/repeats/grader.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • figure: Figure 9.2.4.4.A, printed page 133 (Exact point labels. Protocol Sonnet 4/4.5 labels are represented in the Claude for Life Sciences run to avoid duplicate result rows.) — supports /results
LAB-Bench SeqQA1 run

Open benchmark record →

Evaluation run

lab-bench-seqqa-anthropic-sonnet45-system-card

From Claude Sonnet 4.5 System Card

lab-bench-seqqa-anthropic-10shot-no-toolsvNot reported

Evaluated models / systems: Claude Opus 4, Claude Opus 4.1, Claude Sonnet 4, Claude Sonnet 4.5

Scopetrack
Shots10
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
Statisticsreported point estimate with plotted error bars
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
LAB-Bench scoreabsoluteproportionNot reportedNot reported

Results

ModelMetricValuen
Claude Sonnet 4LAB-Bench score0.682 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported
Claude Opus 4LAB-Bench score0.723 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported
Claude Opus 4.1LAB-Bench score0.785 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported
Claude Sonnet 4.5LAB-Bench score0.78 proportion
Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported.
Not reported

Evidence

  • section: Section 9.2.4.4, printed pages 132–133 (Four selected tracks, 10-shot prompting, no search/bioinformatics tools, threshold, and unreported scope/repeats/grader.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • figure: Figure 9.2.4.4.A, printed page 133 (Exact point labels. Protocol Sonnet 4/4.5 labels are represented in the Claude for Life Sciences run to avoid duplicate result rows.) — supports /results