Exact model identity

Claude Sonnet 4.6

Anthropic · version status: reported

Similar model names are never merged automatically. Evaluation membership and numeric result rows only reference the exact ID claude-sonnet-4-6.

Evaluation settings

Numeric results

BenchmarkWorkRun / groupMetricValue
BioMysteryBenchEvaluating Claude's bioinformatics research capabilities with BioMysteryBenchbiomysterybench-v8-human-difficult
biomysterybench-v8-human-difficult-five-episodes
Accuracy — human-difficult19.1 percent
BioMysteryBenchEvaluating Claude's bioinformatics research capabilities with BioMysteryBenchbiomysterybench-v8-human-solvable
biomysterybench-v8-human-solvable-five-episodes
Accuracy — human-solvable71.8 percent
LAB-Bench FigQAClaude Sonnet 4.6 System Cardlab-bench-figqa-crop-tool
lab-bench-figqa-sonnet46-adaptive-max-crop-tool-five-runs
FigQA score77.1 percent
LAB-Bench FigQAClaude Sonnet 4.6 System Cardlab-bench-figqa-no-tools
lab-bench-figqa-sonnet46-adaptive-max-no-tools-five-runs
FigQA score58.8 percent
SpatialBenchSpatialBench 159-evaluation repository snapshotspatialbench-repo-159-mini-swe-agent
spatialbench-repo-159-mini-swe-agent
Accuracy44.23 percent
SpatialBenchSpatialBench 159-evaluation repository snapshotspatialbench-repo-159-mini-swe-agent
spatialbench-repo-159-mini-swe-agent
Cost0.273 USD per evaluation
SpatialBenchSpatialBench 159-evaluation repository snapshotspatialbench-repo-159-mini-swe-agent
spatialbench-repo-159-mini-swe-agent
Duration405.3 seconds per evaluation