official-release · official model provider

Advancing Claude in healthcare and the life sciences

Anthropic · 2026-01-11

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

evaluation

anthropic-computational-biology-evaluation

Internalunknown

Benchmark: Anthropic Computational Biology Eval · version reported-2026-01-11

Selection
not reported
Metrics
Accuracy delta

Not reported / unresolved: task count; task data; grader; prompt; tools; repeats; absolute accuracy

Only the exact Opus 4.5 improvement annotation is normalized; plotted absolute values are not estimated.

Evidence
  • figure: Evals for key life sciences tasks — Computational biology (Shows Opus 4.1, Sonnet 4.5, Opus 4.5 and an exact +10.5% annotation.)
    Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /reporting_gaps, /notes

benchmark creation

anthropic-key-life-sciences-creation

Internalunknown

Benchmark: Anthropic Key Life Sciences Evals · version reported-2026-01-11

Selection
not applicable
Models
Not reported / not applicable
Metrics
Not reported / not applicable
Linked runs
None

Anthropic is the creator and evaluator of this private internal suite; child-track evaluation relations are recorded separately.

Evidence
  • figure: Evals for key life sciences tasks (Anthropic page presents the three private internal directions and model trend lines.)
    Supports: /benchmark_version, /relation_type, /status, /scope, /notes

evaluation

anthropic-protein-understanding-evaluation

Internalunknown

Benchmark: Anthropic Protein Understanding Eval · version reported-2026-01-11

Selection
not reported
Metrics
Accuracy delta

Not reported / unresolved: task count; task data; grader; prompt; tools; repeats; absolute accuracy

Only the exact Opus 4.5 improvement annotation is normalized; plotted absolute values are not estimated.

Evidence
  • figure: Evals for key life sciences tasks — Protein understanding (Shows Opus 4.1, Sonnet 4.5, Opus 4.5 and an exact +10.3% annotation.)
    Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /reporting_gaps, /notes

evaluation

anthropic-scientific-figure-evaluation

Internalunknown

Benchmark: Anthropic Scientific Figure Interpretation Eval · version reported-2026-01-11

Selection
not reported
Metrics
Accuracy delta

Not reported / unresolved: task count; task data; grader; prompt; tools; repeats; absolute accuracy

Only the exact Opus 4.5 improvement annotation is normalized; plotted absolute values are not estimated.

Evidence
  • figure: Evals for key life sciences tasks — Scientific figure interpretation (Shows Opus 4.1, Sonnet 4.5, Opus 4.5 and an exact +13.2% annotation.)
    Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /reporting_gaps, /notes

external result summary

anthropic-spatialbench-external-summary

External summaryfull · n=146

Benchmark: SpatialBench · version paper-v2

Selection
not applicable · all 146 paper-v2 problems described in the chart
Metrics
Accuracy
Linked runs
None

Not reported / unresolved: Anthropic did not conduct or claim an independent rerun; harness settings are inherited from the cited LatchBio source rather than reported as an Anthropic protocol

The exact rounded chart values are explicitly attributed to LatchBio SpatialBench. This relation is a third-party result summary and is never counted as an Anthropic self-evaluation.

Evidence
  • figure: SpatialBench: Spatial biology analysis by LatchBio (Caption says Source: LatchBio SpatialBench and 146 verifiable problems across five platforms and seven task categories.)
    Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /reporting_gaps, /notes

Normalized evaluation runs

Anthropic Computational Biology Eval1 run

Open benchmark record →

Evaluation run

anthropic-computational-biology-delta

From Advancing Claude in healthcare and the life sciences

anthropic-computational-biology-deltavreported-2026-01-11

Evaluated models / systems: Claude Opus 4.1, Claude Opus 4.5, Claude Sonnet 4.5

Scopeunknown
ShotsNot reported
TurnsNot reported
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNot reported
InternetNot reported
DatabasesNot reported
Code executionNot reported
ContainerNot reported
External toolsNot reported
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
StatisticsNot reported
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5delta
baseline: Claude Opus 4.1
percent delta as annotatedNot reportedNot reported

Results

ModelMetricValuen
Claude Opus 4.5Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5Δ 10.5 percent delta as annotated
Exact +10.5% chart annotation; no absolute accuracy or relative-versus-percentage-point interpretation is inferred.
Not reported

Evidence

  • figure: Evals for key life sciences tasks — Computational biology (Labels Accuracy (%) and annotates +10.5% from Opus 4.1 to Opus 4.5.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics, /results
Anthropic Protein Understanding Eval1 run

Open benchmark record →

Evaluation run

anthropic-protein-understanding-delta

From Advancing Claude in healthcare and the life sciences

anthropic-protein-understanding-deltavreported-2026-01-11

Evaluated models / systems: Claude Opus 4.1, Claude Opus 4.5, Claude Sonnet 4.5

Scopeunknown
ShotsNot reported
TurnsNot reported
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNot reported
InternetNot reported
DatabasesNot reported
Code executionNot reported
ContainerNot reported
External toolsNot reported
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
StatisticsNot reported
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5delta
baseline: Claude Opus 4.1
percent delta as annotatedNot reportedNot reported

Results

ModelMetricValuen
Claude Opus 4.5Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5Δ 10.3 percent delta as annotated
Exact +10.3% chart annotation; no absolute accuracy or relative-versus-percentage-point interpretation is inferred.
Not reported

Evidence

  • figure: Evals for key life sciences tasks — Protein understanding (Labels Accuracy (%) and annotates +10.3% from Opus 4.1 to Opus 4.5.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics, /results
Anthropic Scientific Figure Interpretation Eval1 run

Open benchmark record →

Evaluation run

anthropic-scientific-figure-delta

From Advancing Claude in healthcare and the life sciences

anthropic-scientific-figure-deltavreported-2026-01-11

Evaluated models / systems: Claude Opus 4.1, Claude Opus 4.5, Claude Sonnet 4.5

Scopeunknown
ShotsNot reported
TurnsNot reported
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNot reported
InternetNot reported
DatabasesNot reported
Code executionNot reported
ContainerNot reported
External toolsNot reported
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
StatisticsNot reported
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5delta
baseline: Claude Opus 4.1
percent delta as annotatedNot reportedNot reported

Results

ModelMetricValuen
Claude Opus 4.5Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5Δ 13.2 percent delta as annotated
Exact +13.2% chart annotation; no absolute accuracy or relative-versus-percentage-point interpretation is inferred.
Not reported

Evidence

  • figure: Evals for key life sciences tasks — Scientific figure interpretation (Labels Accuracy (%) and annotates +13.2% from Opus 4.1 to Opus 4.5.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics, /results