official-release · official model provider

Claude for Life Sciences

Anthropic · 2025-10-20

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

evaluation

anthropic-life-sciences-bixbench

Partialunknown

Benchmark: BixBench · version Not reported

Selection
not reported
Metrics
Not reported / not applicable
Linked runs
None

Not reported / unresolved: benchmark version; full, subset, or track scope; realized n; metric and aggregation; numeric results; prompt, tools, and budget; repeats and grader

Anthropic states that Sonnet 4.5 shows a similar improvement over Sonnet 4 on BixBench, but publishes no score or sufficiently specified evaluation setting. No value is inferred from the wording.

Evidence
  • section: Making Claude a better research partner, paragraph 2 (Names BixBench and compares Sonnet 4.5 with predecessor Sonnet 4, without settings, metric, or results.)
    Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /reporting_gaps, /notes

Normalized evaluation runs

LAB-Bench ProtocolQA1 run

Open benchmark record →

Evaluation run

lab-bench-protocolqa-anthropic

From Claude for Life Sciences

lab-bench-protocolqa-anthropic-life-sciences-10shot-no-toolsvNot reported

Evaluated models / systems: Claude Sonnet 4, Claude Sonnet 4.5

Scopetrack
Shots10
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
Statisticsreported point estimate with plotted error bars
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
ProtocolQA scoreabsoluteproportionNot reportedNot reported

Results

ModelMetricValuen
Claude Sonnet 4.5ProtocolQA score0.83 proportion
Official release rounds the system-card point label 0.833 to 0.83.
Not reported
Claude Sonnet 4ProtocolQA score0.74 proportion
Official release rounds the system-card point label 0.741 to 0.74.
Not reported

Evidence

  • section: Section 9.2.4.4 and Figure 9.2.4.4.A, printed pages 132–133 (10-shot, no search/bioinformatics tools, point estimates with error bars, and unreported n/repeats/grader.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • section: Making Claude a better research partner; footnote 1 (Sonnet 4.5 score 0.83, Sonnet 4 score 0.74, and 10-shot prompting.) — supports /results