suite · audited · verified 2026-07-22

Anthropic Key Life Sciences Evals

An Anthropic private internal suite reported only through an official accuracy chart covering scientific figure interpretation, computational biology, and protein understanding.

Private/internal benchmark: tasks, counts, and core protocol are not public. This record is excluded from public and runnable coverage statistics. Exact labeled deltas remain inspectable, but are not converted into absolute scores or plotted as a ranking.

Benchmark definition

What is counted

Version
reported-2026-01-11
Total
Not reported (internal evaluation tasks shown only as three directions in an official chart)
Task formats
private internal accuracy evaluation; exact format not reported
Capabilities
KnowledgeEvidence synthesisData analysisScientific reasoning
Modalities
TextFigure

Version history

VersionStatusRelease / as-ofTotalFormal tracks
reported-2026-01-11
anthropic-key-life-sciences-reported-2026-01-11
current2026-01-11Not reported (internal evaluation tasks shown only as three directions in an official chart)anthropic-scientific-figure-interpretation, anthropic-computational-biology, anthropic-protein-understanding

Registered child tracks

Anthropic Computational Biology Eval

Private Anthropic computational-biology evaluation direction reported only through a model-trend chart, without public tasks, counts, or protocol details.

1 evaluation run(s)

Anthropic Protein Understanding Eval

Private Anthropic protein-understanding evaluation direction reported only through a model-trend chart, without public tasks, counts, or protocol details.

1 evaluation run(s)

Scientific Task Atlas

Scientific task classification

partial for reported-2026-01-11 · as of 2026-01-11. The official chart names three private directions but does not disclose tasks or an exhaustive scientific taxonomy.

The source names only a broad direction; no more specific leaf task can be assigned without inference.

Scientific coverage notes

DomainCoverageCountInterpretation
Life scienceexplicitly-in-scopeNot reportedThree named internal directions are public, but tasks and counts are not.

Relationship registry

How works use this benchmark

Partial claims, non-evaluation uses, and third-party summaries stay visible without entering model comparisons.

Normalized evaluations

evaluation

anthropic-computational-biology-evaluation

Internalunknown

Benchmark: Anthropic Computational Biology Eval · version reported-2026-01-11

Work: Advancing Claude in healthcare and the life sciences · source version anthropic-healthcare-life-sciences-2026-01-11

Selection
not reported
Metrics
Accuracy delta

Not reported / unresolved: task count; task data; grader; prompt; tools; repeats; absolute accuracy

Only the exact Opus 4.5 improvement annotation is normalized; plotted absolute values are not estimated.

Evidence
  • figure: Evals for key life sciences tasks — Computational biology (Shows Opus 4.1, Sonnet 4.5, Opus 4.5 and an exact +10.5% annotation.)
    Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /reporting_gaps, /notes

evaluation

anthropic-protein-understanding-evaluation

Internalunknown

Benchmark: Anthropic Protein Understanding Eval · version reported-2026-01-11

Work: Advancing Claude in healthcare and the life sciences · source version anthropic-healthcare-life-sciences-2026-01-11

Selection
not reported
Metrics
Accuracy delta

Not reported / unresolved: task count; task data; grader; prompt; tools; repeats; absolute accuracy

Only the exact Opus 4.5 improvement annotation is normalized; plotted absolute values are not estimated.

Evidence
  • figure: Evals for key life sciences tasks — Protein understanding (Shows Opus 4.1, Sonnet 4.5, Opus 4.5 and an exact +10.3% annotation.)
    Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /reporting_gaps, /notes

evaluation

anthropic-scientific-figure-evaluation

Internalunknown

Benchmark: Anthropic Scientific Figure Interpretation Eval · version reported-2026-01-11

Work: Advancing Claude in healthcare and the life sciences · source version anthropic-healthcare-life-sciences-2026-01-11

Selection
not reported
Metrics
Accuracy delta

Not reported / unresolved: task count; task data; grader; prompt; tools; repeats; absolute accuracy

Only the exact Opus 4.5 improvement annotation is normalized; plotted absolute values are not estimated.

Evidence
  • figure: Evals for key life sciences tasks — Scientific figure interpretation (Shows Opus 4.1, Sonnet 4.5, Opus 4.5 and an exact +13.2% annotation.)
    Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /reporting_gaps, /notes

Creation, training, validation, or model-selection uses

benchmark creation

anthropic-key-life-sciences-creation

Internalunknown

Work: Advancing Claude in healthcare and the life sciences · source version anthropic-healthcare-life-sciences-2026-01-11

Selection
not applicable
Models
Not reported / not applicable
Metrics
Not reported / not applicable
Linked runs
None

Anthropic is the creator and evaluator of this private internal suite; child-track evaluation relations are recorded separately.

Evidence
  • figure: Evals for key life sciences tasks (Anthropic page presents the three private internal directions and model trend lines.)
    Supports: /benchmark_version, /relation_type, /status, /scope, /notes

Evaluation registry

Works and run settings

A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.

Evaluation run

anthropic-computational-biology-delta

From Advancing Claude in healthcare and the life sciences

anthropic-computational-biology-deltavreported-2026-01-11

Evaluated models / systems: Claude Opus 4.1, Claude Opus 4.5, Claude Sonnet 4.5

Scopeunknown
ShotsNot reported
TurnsNot reported
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNot reported
InternetNot reported
DatabasesNot reported
Code executionNot reported
ContainerNot reported
External toolsNot reported
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
StatisticsNot reported
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5delta
baseline: Claude Opus 4.1
percent delta as annotatedNot reportedNot reported

Results

ModelMetricValuen
Claude Opus 4.5Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5Δ 10.5 percent delta as annotated
Exact +10.5% chart annotation; no absolute accuracy or relative-versus-percentage-point interpretation is inferred.
Not reported

Evidence

  • figure: Evals for key life sciences tasks — Computational biology (Labels Accuracy (%) and annotates +10.5% from Opus 4.1 to Opus 4.5.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics, /results

Evaluation run

anthropic-protein-understanding-delta

From Advancing Claude in healthcare and the life sciences

anthropic-protein-understanding-deltavreported-2026-01-11

Evaluated models / systems: Claude Opus 4.1, Claude Opus 4.5, Claude Sonnet 4.5

Scopeunknown
ShotsNot reported
TurnsNot reported
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNot reported
InternetNot reported
DatabasesNot reported
Code executionNot reported
ContainerNot reported
External toolsNot reported
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
StatisticsNot reported
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5delta
baseline: Claude Opus 4.1
percent delta as annotatedNot reportedNot reported

Results

ModelMetricValuen
Claude Opus 4.5Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5Δ 10.3 percent delta as annotated
Exact +10.3% chart annotation; no absolute accuracy or relative-versus-percentage-point interpretation is inferred.
Not reported

Evidence

  • figure: Evals for key life sciences tasks — Protein understanding (Labels Accuracy (%) and annotates +10.3% from Opus 4.1 to Opus 4.5.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics, /results

Evaluation run

anthropic-scientific-figure-delta

From Advancing Claude in healthcare and the life sciences

anthropic-scientific-figure-deltavreported-2026-01-11

Evaluated models / systems: Claude Opus 4.1, Claude Opus 4.5, Claude Sonnet 4.5

Scopeunknown
ShotsNot reported
TurnsNot reported
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNot reported
InternetNot reported
DatabasesNot reported
Code executionNot reported
ContainerNot reported
External toolsNot reported
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
StatisticsNot reported
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5delta
baseline: Claude Opus 4.1
percent delta as annotatedNot reportedNot reported

Results

ModelMetricValuen
Claude Opus 4.5Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5Δ 13.2 percent delta as annotated
Exact +13.2% chart annotation; no absolute accuracy or relative-versus-percentage-point interpretation is inferred.
Not reported

Evidence

  • figure: Evals for key life sciences tasks — Scientific figure interpretation (Labels Accuracy (%) and annotates +13.2% from Opus 4.1 to Opus 4.5.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics, /results

Evidence and change history

Source locators remain visible; expand an item to inspect the exact Registry fields it supports.

Advancing Claude in healthcare and the life sciences · figure: Evals for key life sciences tasks (Names the three internal evaluation directions, labels Accuracy (%), and annotates Opus 4.5 changes from Opus 4.1.) · Supports 30 fields

Open source →

  • /name
  • /aliases
  • /summary
  • /kind
  • /parent_id
  • /organizations
  • /release_date
  • /latest_version
  • /domains
  • /capabilities
  • /modalities
  • /task_formats
  • /task_counts/total
  • /task_counts/basis
  • /task_counts/subsets
  • /coverage_notes
  • /access/level
  • /access/tasks
  • /access/artifacts
  • /access/grader
  • /access/license
  • /access/biosafety_notes
  • /resources
  • /implementations
  • /versions/0/release_date
  • /versions/0/as_of
  • /versions/0/task_counts/total
  • /versions/0/task_counts/basis
  • /versions/0/task_counts/subsets
  • /versions/0/formal_tracks

View source-level modification history on GitHub →