system-card · official model provider

Claude Sonnet 4.6 System Card

Anthropic · 2026-02-17

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

No BenchmarkUse relation is normalized for this legacy work yet. Existing EvaluationRuns remain available below.

Normalized evaluation runs

LAB-Bench FigQA2 runs

Open benchmark record →

Evaluation run

lab-bench-figqa-crop-tool

From Claude Sonnet 4.6 System Card

lab-bench-figqa-sonnet46-adaptive-max-crop-tool-five-runsvNot reported

Evaluated models / systems: Claude Opus 4.6, Claude Sonnet 4.5, Claude Sonnet 4.6

Scopetrack
ShotsNot reported
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortadaptive thinking at max effort
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolssimple image cropping tool
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats5
GraderNot reported
Statisticsmean over five runs with 95% CI shown
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
FigQA scoreabsolutepercentmean over five runsNot reported

Results

ModelMetricValuen
Claude Sonnet 4.6FigQA score77.1 percent
95% CI is plotted but numeric bounds are not reported.
Not reported
Claude Sonnet 4.5FigQA score59.3 percent
95% CI is plotted but numeric bounds are not reported.
Not reported
Claude Opus 4.6FigQA score78.3 percent
95% CI is plotted but numeric bounds are not reported.
Not reported

Evidence

  • section: Section 2.17.1, printed page 33 (Adaptive thinking, max effort, simple crop tool, five runs, 95% CI, and unreported n/snapshot/grader.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • section: Section 2.17.1 and Figure 2.17.1.A, printed page 33 (All three crop-tool point estimates.) — supports /results

Evaluation run

lab-bench-figqa-no-tools

From Claude Sonnet 4.6 System Card

lab-bench-figqa-sonnet46-adaptive-max-no-tools-five-runsvNot reported

Evaluated models / systems: Claude Opus 4.6, Claude Sonnet 4.5, Claude Sonnet 4.6

Scopetrack
ShotsNot reported
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortadaptive thinking at max effort
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats5
GraderNot reported
Statisticsmean over five runs with 95% CI shown
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
FigQA scoreabsolutepercentmean over five runsNot reported

Results

ModelMetricValuen
Claude Sonnet 4.6FigQA score58.8 percent
95% CI is plotted but numeric bounds are not reported.
Not reported
Claude Sonnet 4.5FigQA score53.4 percent
95% CI is plotted but numeric bounds are not reported.
Not reported
Claude Opus 4.6FigQA score58 percent
95% CI is plotted but numeric bounds are not reported.
Not reported

Evidence

  • section: Section 2.17.1, printed page 33 (Adaptive thinking, max effort, no tools, five runs, 95% CI, and unreported n/snapshot/grader.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • section: Section 2.17.1 and Figure 2.17.1.A, printed page 33 (All three no-tool point estimates.) — supports /results