system-card · official model provider
Claude Sonnet 4.5 System Card
Anthropic · 2025-09-29
Relationship layer
Benchmark usage
This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.
No BenchmarkUse relation is normalized for this legacy work yet. Existing EvaluationRuns remain available below.
Normalized evaluation runs
LAB-Bench CloningScenarios1 run
Open benchmark record →
lab-bench-cloning-scenarios-anthropic-10shot-no-toolsvNot reported
Evaluated models / systems: Claude Opus 4, Claude Opus 4.1, Claude Sonnet 4, Claude Sonnet 4.5
Scopetrack
Shots10
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
Statisticsreported point estimate with plotted error bars
ContaminationNot reported
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| LAB-Bench score | absolute | proportion | Not reported | Not reported |
Results
| Model | Metric | Value | n |
|---|
| Claude Sonnet 4 | LAB-Bench score | 0.485 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
| Claude Opus 4 | LAB-Bench score | 0.545 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
| Claude Opus 4.1 | LAB-Bench score | 0.758 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
| Claude Sonnet 4.5 | LAB-Bench score | 0.667 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
Evidence
- section: Section 9.2.4.4, printed pages 132–133 (Four selected tracks, 10-shot prompting, no search/bioinformatics tools, threshold, and unreported scope/repeats/grader.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
- figure: Figure 9.2.4.4.A, printed page 133 (Exact point labels. Protocol Sonnet 4/4.5 labels are represented in the Claude for Life Sciences run to avoid duplicate result rows.) — supports /results
LAB-Bench FigQA1 run
Open benchmark record →
lab-bench-figqa-anthropic-10shot-no-toolsvNot reported
Evaluated models / systems: Claude Opus 4, Claude Opus 4.1, Claude Sonnet 4, Claude Sonnet 4.5
Scopetrack
Shots10
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
Statisticsreported point estimate with plotted error bars
ContaminationNot reported
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| LAB-Bench score | absolute | proportion | Not reported | Not reported |
Results
| Model | Metric | Value | n |
|---|
| Claude Sonnet 4 | LAB-Bench score | 0.398 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
| Claude Opus 4 | LAB-Bench score | 0.508 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
| Claude Opus 4.1 | LAB-Bench score | 0.481 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
| Claude Sonnet 4.5 | LAB-Bench score | 0.497 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
Evidence
- section: Section 9.2.4.4, printed pages 132–133 (Four selected tracks, 10-shot prompting, no search/bioinformatics tools, threshold, and unreported scope/repeats/grader.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
- figure: Figure 9.2.4.4.A, printed page 133 (Exact point labels. Protocol Sonnet 4/4.5 labels are represented in the Claude for Life Sciences run to avoid duplicate result rows.) — supports /results
LAB-Bench ProtocolQA1 run
Open benchmark record →
lab-bench-protocolqa-anthropic-10shot-no-toolsvNot reported
Evaluated models / systems: Claude Opus 4, Claude Opus 4.1
Scopetrack
Shots10
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
Statisticsreported point estimate with plotted error bars
ContaminationNot reported
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| LAB-Bench score | absolute | proportion | Not reported | Not reported |
Results
| Model | Metric | Value | n |
|---|
| Claude Opus 4 | LAB-Bench score | 0.796 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
| Claude Opus 4.1 | LAB-Bench score | 0.833 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
Evidence
- section: Section 9.2.4.4, printed pages 132–133 (Four selected tracks, 10-shot prompting, no search/bioinformatics tools, threshold, and unreported scope/repeats/grader.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
- figure: Figure 9.2.4.4.A, printed page 133 (Exact point labels. Protocol Sonnet 4/4.5 labels are represented in the Claude for Life Sciences run to avoid duplicate result rows.) — supports /results
LAB-Bench SeqQA1 run
Open benchmark record →
lab-bench-seqqa-anthropic-10shot-no-toolsvNot reported
Evaluated models / systems: Claude Opus 4, Claude Opus 4.1, Claude Sonnet 4, Claude Sonnet 4.5
Scopetrack
Shots10
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
Statisticsreported point estimate with plotted error bars
ContaminationNot reported
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| LAB-Bench score | absolute | proportion | Not reported | Not reported |
Results
| Model | Metric | Value | n |
|---|
| Claude Sonnet 4 | LAB-Bench score | 0.682 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
| Claude Opus 4 | LAB-Bench score | 0.723 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
| Claude Opus 4.1 | LAB-Bench score | 0.785 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
| Claude Sonnet 4.5 | LAB-Bench score | 0.78 proportion Point label transcribed from Figure 9.2.4.4.A; numeric error-bar bounds are not reported. | Not reported |
Evidence
- section: Section 9.2.4.4, printed pages 132–133 (Four selected tracks, 10-shot prompting, no search/bioinformatics tools, threshold, and unreported scope/repeats/grader.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
- figure: Figure 9.2.4.4.A, printed page 133 (Exact point labels. Protocol Sonnet 4/4.5 labels are represented in the Claude for Life Sciences run to avoid duplicate result rows.) — supports /results