evaluation
anthropic-life-sciences-bixbench
Work: Claude for Life Sciences · source version anthropic-life-sciences-2025-10-20
- Selection
- not reported
- Models
- Claude Sonnet 4.5, Claude Sonnet 4
- Metrics
- Not reported / not applicable
- Linked runs
- None
Not reported / unresolved: benchmark version; full, subset, or track scope; realized n; metric and aggregation; numeric results; prompt, tools, and budget; repeats and grader
Anthropic states that Sonnet 4.5 shows a similar improvement over Sonnet 4 on BixBench, but publishes no score or sufficiently specified evaluation setting. No value is inferred from the wording.
Evidence
- section: Making Claude a better research partner, paragraph 2 (Names BixBench and compares Sonnet 4.5 with predecessor Sonnet 4, without settings, metric, or results.)
Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /reporting_gaps, /notes