track · audited · verified 2026-07-22

Anthropic Scientific Figure Interpretation Eval

Private Anthropic evaluation direction for scientific figure interpretation, reported only through a model-trend chart with no task count or released examples.

Private/internal benchmark: tasks, counts, and core protocol are not public. This record is excluded from public and runnable coverage statistics. Exact labeled deltas remain inspectable, but are not converted into absolute scores or plotted as a ranking.

Benchmark definition

What is counted

Version
reported-2026-01-11
Total
Not reported (private scientific figure interpretation tasks)
Task formats
private scientific figure interpretation; exact format not reported
Capabilities
Evidence synthesisScientific reasoning
Modalities
TextFigure

Version history

VersionStatusRelease / as-ofTotalFormal tracks
reported-2026-01-11
anthropic-scientific-figure-reported-2026-01-11
current2026-01-11Not reported (private scientific figure interpretation tasks)None registered

Scientific Task Atlas

Scientific task classification

complete for reported-2026-01-11 · as of 2026-01-11. The public direction label maps directly to scientific evidence interpretation; its private task count is not reported.

Scientific taskCoverageCountMappingEvidence
Scientific evidence interpretationexplicitly-in-scopeNot reported
private scientific figure interpretation tasks
official-track
high confidence
anthropic-scientific-figure-evidence
No task-level examples or count are public.

Scientific coverage notes

DomainCoverageCountInterpretation
Life scienceexplicitly-in-scopeNot reportedThe direction is explicit; task count is not reported.

Relationship registry

How works use this benchmark

Partial claims, non-evaluation uses, and third-party summaries stay visible without entering model comparisons.

Normalized evaluations

evaluation

anthropic-scientific-figure-evaluation

Internalunknown

Work: Advancing Claude in healthcare and the life sciences · source version anthropic-healthcare-life-sciences-2026-01-11

Selection
not reported
Metrics
Accuracy delta

Not reported / unresolved: task count; task data; grader; prompt; tools; repeats; absolute accuracy

Only the exact Opus 4.5 improvement annotation is normalized; plotted absolute values are not estimated.

Evidence
  • figure: Evals for key life sciences tasks — Scientific figure interpretation (Shows Opus 4.1, Sonnet 4.5, Opus 4.5 and an exact +13.2% annotation.)
    Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /reporting_gaps, /notes

Evaluation registry

Works and run settings

A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.

Evaluation run

anthropic-scientific-figure-delta

From Advancing Claude in healthcare and the life sciences

anthropic-scientific-figure-deltavreported-2026-01-11

Evaluated models / systems: Claude Opus 4.1, Claude Opus 4.5, Claude Sonnet 4.5

Scopeunknown
ShotsNot reported
TurnsNot reported
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNot reported
InternetNot reported
DatabasesNot reported
Code executionNot reported
ContainerNot reported
External toolsNot reported
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
StatisticsNot reported
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5delta
baseline: Claude Opus 4.1
percent delta as annotatedNot reportedNot reported

Results

ModelMetricValuen
Claude Opus 4.5Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5Δ 13.2 percent delta as annotated
Exact +13.2% chart annotation; no absolute accuracy or relative-versus-percentage-point interpretation is inferred.
Not reported

Evidence

  • figure: Evals for key life sciences tasks — Scientific figure interpretation (Labels Accuracy (%) and annotates +13.2% from Opus 4.1 to Opus 4.5.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics, /results

Evidence and change history

Source locators remain visible; expand an item to inspect the exact Registry fields it supports.

Advancing Claude in healthcare and the life sciences · figure: Evals for key life sciences tasks — Scientific figure interpretation (Exact direction label and +13.2% Opus 4.5 annotation.) · Supports 29 fields

Open source →

  • /name
  • /summary
  • /kind
  • /parent_id
  • /organizations
  • /release_date
  • /latest_version
  • /domains
  • /capabilities
  • /modalities
  • /task_formats
  • /task_counts/total
  • /task_counts/basis
  • /task_counts/subsets
  • /coverage_notes
  • /access/level
  • /access/tasks
  • /access/artifacts
  • /access/grader
  • /access/license
  • /access/biosafety_notes
  • /resources
  • /implementations
  • /versions/0/release_date
  • /versions/0/as_of
  • /versions/0/task_counts/total
  • /versions/0/task_counts/basis
  • /versions/0/task_counts/subsets
  • /scientific_task_classification/entries/0

View source-level modification history on GitHub →