official-release · official model provider
Advancing Claude in healthcare and the life sciences
Anthropic · 2026-01-11
Relationship layer
Benchmark usage
This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.
evaluation
anthropic-computational-biology-evaluation
Internalunknown
Benchmark: Anthropic Computational Biology Eval · version reported-2026-01-11
- Selection
- not reported
- Metrics
- Accuracy delta
Not reported / unresolved: task count; task data; grader; prompt; tools; repeats; absolute accuracy
Only the exact Opus 4.5 improvement annotation is normalized; plotted absolute values are not estimated.
Evidence
- figure: Evals for key life sciences tasks — Computational biology (Shows Opus 4.1, Sonnet 4.5, Opus 4.5 and an exact +10.5% annotation.)
Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /reporting_gaps, /notes
benchmark creation
anthropic-key-life-sciences-creation
Internalunknown
Benchmark: Anthropic Key Life Sciences Evals · version reported-2026-01-11
- Selection
- not applicable
- Models
- Not reported / not applicable
- Metrics
- Not reported / not applicable
- Linked runs
- None
Anthropic is the creator and evaluator of this private internal suite; child-track evaluation relations are recorded separately.
Evidence
- figure: Evals for key life sciences tasks (Anthropic page presents the three private internal directions and model trend lines.)
Supports: /benchmark_version, /relation_type, /status, /scope, /notes
evaluation
anthropic-protein-understanding-evaluation
Internalunknown
Benchmark: Anthropic Protein Understanding Eval · version reported-2026-01-11
- Selection
- not reported
- Metrics
- Accuracy delta
Not reported / unresolved: task count; task data; grader; prompt; tools; repeats; absolute accuracy
Only the exact Opus 4.5 improvement annotation is normalized; plotted absolute values are not estimated.
Evidence
- figure: Evals for key life sciences tasks — Protein understanding (Shows Opus 4.1, Sonnet 4.5, Opus 4.5 and an exact +10.3% annotation.)
Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /reporting_gaps, /notes
evaluation
anthropic-scientific-figure-evaluation
Internalunknown
Benchmark: Anthropic Scientific Figure Interpretation Eval · version reported-2026-01-11
- Selection
- not reported
- Metrics
- Accuracy delta
Not reported / unresolved: task count; task data; grader; prompt; tools; repeats; absolute accuracy
Only the exact Opus 4.5 improvement annotation is normalized; plotted absolute values are not estimated.
Evidence
- figure: Evals for key life sciences tasks — Scientific figure interpretation (Shows Opus 4.1, Sonnet 4.5, Opus 4.5 and an exact +13.2% annotation.)
Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /reporting_gaps, /notes
external result summary
anthropic-spatialbench-external-summary
External summaryfull · n=146
Benchmark: SpatialBench · version paper-v2
- Selection
- not applicable · all 146 paper-v2 problems described in the chart
- Metrics
- Accuracy
- Linked runs
- None
Not reported / unresolved: Anthropic did not conduct or claim an independent rerun; harness settings are inherited from the cited LatchBio source rather than reported as an Anthropic protocol
The exact rounded chart values are explicitly attributed to LatchBio SpatialBench. This relation is a third-party result summary and is never counted as an Anthropic self-evaluation.
Evidence
- figure: SpatialBench: Spatial biology analysis by LatchBio (Caption says Source: LatchBio SpatialBench and 146 verifiable problems across five platforms and seven task categories.)
Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /reporting_gaps, /notes
Normalized evaluation runs
Anthropic Computational Biology Eval1 run
Open benchmark record →
anthropic-computational-biology-deltavreported-2026-01-11
Evaluated models / systems: Claude Opus 4.1, Claude Opus 4.5, Claude Sonnet 4.5
Scopeunknown
ShotsNot reported
TurnsNot reported
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNot reported
InternetNot reported
DatabasesNot reported
Code executionNot reported
ContainerNot reported
External toolsNot reported
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
StatisticsNot reported
ContaminationNot reported
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5 | delta baseline: Claude Opus 4.1 | percent delta as annotated | Not reported | Not reported |
Results
| Model | Metric | Value | n |
|---|
| Claude Opus 4.5 | Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5 | Δ 10.5 percent delta as annotated Exact +10.5% chart annotation; no absolute accuracy or relative-versus-percentage-point interpretation is inferred. | Not reported |
Evidence
- figure: Evals for key life sciences tasks — Computational biology (Labels Accuracy (%) and annotates +10.5% from Opus 4.1 to Opus 4.5.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics, /results
Anthropic Protein Understanding Eval1 run
Open benchmark record →
anthropic-protein-understanding-deltavreported-2026-01-11
Evaluated models / systems: Claude Opus 4.1, Claude Opus 4.5, Claude Sonnet 4.5
Scopeunknown
ShotsNot reported
TurnsNot reported
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNot reported
InternetNot reported
DatabasesNot reported
Code executionNot reported
ContainerNot reported
External toolsNot reported
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
StatisticsNot reported
ContaminationNot reported
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5 | delta baseline: Claude Opus 4.1 | percent delta as annotated | Not reported | Not reported |
Results
| Model | Metric | Value | n |
|---|
| Claude Opus 4.5 | Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5 | Δ 10.3 percent delta as annotated Exact +10.3% chart annotation; no absolute accuracy or relative-versus-percentage-point interpretation is inferred. | Not reported |
Evidence
- figure: Evals for key life sciences tasks — Protein understanding (Labels Accuracy (%) and annotates +10.3% from Opus 4.1 to Opus 4.5.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics, /results
Anthropic Scientific Figure Interpretation Eval1 run
Open benchmark record →
anthropic-scientific-figure-deltavreported-2026-01-11
Evaluated models / systems: Claude Opus 4.1, Claude Opus 4.5, Claude Sonnet 4.5
Scopeunknown
ShotsNot reported
TurnsNot reported
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNot reported
InternetNot reported
DatabasesNot reported
Code executionNot reported
ContainerNot reported
External toolsNot reported
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
StatisticsNot reported
ContaminationNot reported
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5 | delta baseline: Claude Opus 4.1 | percent delta as annotated | Not reported | Not reported |
Results
| Model | Metric | Value | n |
|---|
| Claude Opus 4.5 | Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5 | Δ 13.2 percent delta as annotated Exact +13.2% chart annotation; no absolute accuracy or relative-versus-percentage-point interpretation is inferred. | Not reported |
Evidence
- figure: Evals for key life sciences tasks — Scientific figure interpretation (Labels Accuracy (%) and annotates +13.2% from Opus 4.1 to Opus 4.5.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics, /results