Exact model identity

Claude Sonnet 4.5

Anthropic · version status: reported

Similar model names are never merged automatically. Evaluation membership and numeric result rows only reference the exact ID claude-sonnet-4-5.

Evaluation settings

BenchmarkWorkRun / groupScope
Anthropic Computational Biology EvalAdvancing Claude in healthcare and the life sciencesanthropic-computational-biology-delta
anthropic-computational-biology-delta
unknown
Anthropic Protein Understanding EvalAdvancing Claude in healthcare and the life sciencesanthropic-protein-understanding-delta
anthropic-protein-understanding-delta
unknown
Anthropic Scientific Figure Interpretation EvalAdvancing Claude in healthcare and the life sciencesanthropic-scientific-figure-delta
anthropic-scientific-figure-delta
unknown
LAB-Bench CloningScenariosClaude Sonnet 4.5 System Cardlab-bench-cloning-scenarios-anthropic-sonnet45-system-card
lab-bench-cloning-scenarios-anthropic-10shot-no-tools
track
LAB-Bench FigQAClaude Sonnet 4.5 System Cardlab-bench-figqa-anthropic-sonnet45-system-card
lab-bench-figqa-anthropic-10shot-no-tools
track
LAB-Bench FigQAClaude Sonnet 4.6 System Cardlab-bench-figqa-crop-tool
lab-bench-figqa-sonnet46-adaptive-max-crop-tool-five-runs
track
LAB-Bench FigQAClaude Sonnet 4.6 System Cardlab-bench-figqa-no-tools
lab-bench-figqa-sonnet46-adaptive-max-no-tools-five-runs
track
LAB-Bench ProtocolQAClaude for Life Scienceslab-bench-protocolqa-anthropic
lab-bench-protocolqa-anthropic-life-sciences-10shot-no-tools
track
LAB-Bench SeqQAClaude Sonnet 4.5 System Cardlab-bench-seqqa-anthropic-sonnet45-system-card
lab-bench-seqqa-anthropic-10shot-no-tools
track
SpatialBenchSpatialBench: Can Agents Analyze Real-World Spatial Biology Data?spatialbench-paper-v2-base
spatialbench-paper-v2-base
full · n=146
SpatialBenchSpatialBench: Can Agents Analyze Real-World Spatial Biology Data?spatialbench-paper-v2-claude-code
spatialbench-paper-v2-claude-code
full · n=146
SpatialBenchSpatialBench 159-evaluation repository snapshotspatialbench-repo-159-mini-swe-agent
spatialbench-repo-159-mini-swe-agent
full · n=159

Numeric results

BenchmarkWorkRun / groupMetricValue
LAB-Bench CloningScenariosClaude Sonnet 4.5 System Cardlab-bench-cloning-scenarios-anthropic-sonnet45-system-card
lab-bench-cloning-scenarios-anthropic-10shot-no-tools
LAB-Bench score0.667 proportion
LAB-Bench FigQAClaude Sonnet 4.5 System Cardlab-bench-figqa-anthropic-sonnet45-system-card
lab-bench-figqa-anthropic-10shot-no-tools
LAB-Bench score0.497 proportion
LAB-Bench FigQAClaude Sonnet 4.6 System Cardlab-bench-figqa-crop-tool
lab-bench-figqa-sonnet46-adaptive-max-crop-tool-five-runs
FigQA score59.3 percent
LAB-Bench FigQAClaude Sonnet 4.6 System Cardlab-bench-figqa-no-tools
lab-bench-figqa-sonnet46-adaptive-max-no-tools-five-runs
FigQA score53.4 percent
LAB-Bench ProtocolQAClaude for Life Scienceslab-bench-protocolqa-anthropic
lab-bench-protocolqa-anthropic-life-sciences-10shot-no-tools
ProtocolQA score0.83 proportion
LAB-Bench SeqQAClaude Sonnet 4.5 System Cardlab-bench-seqqa-anthropic-sonnet45-system-card
lab-bench-seqqa-anthropic-10shot-no-tools
LAB-Bench score0.78 proportion
SpatialBenchSpatialBench: Can Agents Analyze Real-World Spatial Biology Data?spatialbench-paper-v2-base
spatialbench-paper-v2-base
Accuracy28.31 percent
SpatialBenchSpatialBench: Can Agents Analyze Real-World Spatial Biology Data?spatialbench-paper-v2-base
spatialbench-paper-v2-base
Steps2.43 steps per evaluation
SpatialBenchSpatialBench: Can Agents Analyze Real-World Spatial Biology Data?spatialbench-paper-v2-base
spatialbench-paper-v2-base
Latency115.6 seconds per evaluation
SpatialBenchSpatialBench: Can Agents Analyze Real-World Spatial Biology Data?spatialbench-paper-v2-base
spatialbench-paper-v2-base
Cost0.081 USD per evaluation
SpatialBenchSpatialBench: Can Agents Analyze Real-World Spatial Biology Data?spatialbench-paper-v2-claude-code
spatialbench-paper-v2-claude-code
Accuracy45.1 percent
SpatialBenchSpatialBench 159-evaluation repository snapshotspatialbench-repo-159-mini-swe-agent
spatialbench-repo-159-mini-swe-agent
Accuracy41.51 percent
SpatialBenchSpatialBench 159-evaluation repository snapshotspatialbench-repo-159-mini-swe-agent
spatialbench-repo-159-mini-swe-agent
Cost0.2247 USD per evaluation
SpatialBenchSpatialBench 159-evaluation repository snapshotspatialbench-repo-159-mini-swe-agent
spatialbench-repo-159-mini-swe-agent
Duration294.44 seconds per evaluation