system-card · official model provider
Claude Sonnet 4.6 System Card
Anthropic · 2026-02-17
Relationship layer
Benchmark usage
This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.
No BenchmarkUse relation is normalized for this legacy work yet. Existing EvaluationRuns remain available below.
Normalized evaluation runs
LAB-Bench FigQA2 runs
Open benchmark record →
lab-bench-figqa-sonnet46-adaptive-max-crop-tool-five-runsvNot reported
Evaluated models / systems: Claude Opus 4.6, Claude Sonnet 4.5, Claude Sonnet 4.6
Scopetrack
ShotsNot reported
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortadaptive thinking at max effort
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolssimple image cropping tool
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats5
GraderNot reported
Statisticsmean over five runs with 95% CI shown
ContaminationNot reported
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| FigQA score | absolute | percent | mean over five runs | Not reported |
Results
| Model | Metric | Value | n |
|---|
| Claude Sonnet 4.6 | FigQA score | 77.1 percent 95% CI is plotted but numeric bounds are not reported. | Not reported |
| Claude Sonnet 4.5 | FigQA score | 59.3 percent 95% CI is plotted but numeric bounds are not reported. | Not reported |
| Claude Opus 4.6 | FigQA score | 78.3 percent 95% CI is plotted but numeric bounds are not reported. | Not reported |
Evidence
- section: Section 2.17.1, printed page 33 (Adaptive thinking, max effort, simple crop tool, five runs, 95% CI, and unreported n/snapshot/grader.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
- section: Section 2.17.1 and Figure 2.17.1.A, printed page 33 (All three crop-tool point estimates.) — supports /results
lab-bench-figqa-sonnet46-adaptive-max-no-tools-five-runsvNot reported
Evaluated models / systems: Claude Opus 4.6, Claude Sonnet 4.5, Claude Sonnet 4.6
Scopetrack
ShotsNot reported
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortadaptive thinking at max effort
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats5
GraderNot reported
Statisticsmean over five runs with 95% CI shown
ContaminationNot reported
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| FigQA score | absolute | percent | mean over five runs | Not reported |
Results
| Model | Metric | Value | n |
|---|
| Claude Sonnet 4.6 | FigQA score | 58.8 percent 95% CI is plotted but numeric bounds are not reported. | Not reported |
| Claude Sonnet 4.5 | FigQA score | 53.4 percent 95% CI is plotted but numeric bounds are not reported. | Not reported |
| Claude Opus 4.6 | FigQA score | 58 percent 95% CI is plotted but numeric bounds are not reported. | Not reported |
Evidence
- section: Section 2.17.1, printed page 33 (Adaptive thinking, max effort, no tools, five runs, 95% CI, and unreported n/snapshot/grader.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
- section: Section 2.17.1 and Figure 2.17.1.A, printed page 33 (All three no-tool point estimates.) — supports /results