official-release · official model provider
Claude for Life Sciences
Anthropic · 2025-10-20
Relationship layer
Benchmark usage
This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.
evaluation
anthropic-life-sciences-bixbench
Partialunknown
Benchmark: BixBench · version Not reported
- Selection
- not reported
- Metrics
- Not reported / not applicable
- Linked runs
- None
Not reported / unresolved: benchmark version; full, subset, or track scope; realized n; metric and aggregation; numeric results; prompt, tools, and budget; repeats and grader
Anthropic states that Sonnet 4.5 shows a similar improvement over Sonnet 4 on BixBench, but publishes no score or sufficiently specified evaluation setting. No value is inferred from the wording.
Evidence
- section: Making Claude a better research partner, paragraph 2 (Names BixBench and compares Sonnet 4.5 with predecessor Sonnet 4, without settings, metric, or results.)
Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /reporting_gaps, /notes
Normalized evaluation runs
LAB-Bench ProtocolQA1 run
Open benchmark record →
lab-bench-protocolqa-anthropic-life-sciences-10shot-no-toolsvNot reported
Evaluated models / systems: Claude Sonnet 4, Claude Sonnet 4.5
Scopetrack
Shots10
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderNot reported
Statisticsreported point estimate with plotted error bars
ContaminationNot reported
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| ProtocolQA score | absolute | proportion | Not reported | Not reported |
Results
| Model | Metric | Value | n |
|---|
| Claude Sonnet 4.5 | ProtocolQA score | 0.83 proportion Official release rounds the system-card point label 0.833 to 0.83. | Not reported |
| Claude Sonnet 4 | ProtocolQA score | 0.74 proportion Official release rounds the system-card point label 0.741 to 0.74. | Not reported |
Evidence
- section: Section 9.2.4.4 and Figure 9.2.4.4.A, printed pages 132–133 (10-shot, no search/bioinformatics tools, point estimates with error bars, and unreported n/repeats/grader.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
- section: Making Claude a better research partner; footnote 1 (Sonnet 4.5 score 0.83, Sonnet 4 score 0.74, and 10-shot prompting.) — supports /results