Anthropic Computational Biology Eval
Private Anthropic computational-biology evaluation direction reported only through a model-trend chart, without public tasks, counts, or protocol details.
1 evaluation run(s)
suite · audited · verified 2026-07-22
An Anthropic private internal suite reported only through an official accuracy chart covering scientific figure interpretation, computational biology, and protein understanding.
Benchmark definition
| Version | Status | Release / as-of | Total | Formal tracks |
|---|---|---|---|---|
reported-2026-01-11anthropic-key-life-sciences-reported-2026-01-11 | current | 2026-01-11 | Not reported (internal evaluation tasks shown only as three directions in an official chart) | anthropic-scientific-figure-interpretation, anthropic-computational-biology, anthropic-protein-understanding |
Private Anthropic computational-biology evaluation direction reported only through a model-trend chart, without public tasks, counts, or protocol details.
1 evaluation run(s)
Private Anthropic protein-understanding evaluation direction reported only through a model-trend chart, without public tasks, counts, or protocol details.
1 evaluation run(s)
Private Anthropic evaluation direction for scientific figure interpretation, reported only through a model-trend chart with no task count or released examples.
1 evaluation run(s)
Scientific Task Atlas
partial for reported-2026-01-11 · as of 2026-01-11. The official chart names three private directions but does not disclose tasks or an exhaustive scientific taxonomy.
| Domain | Coverage | Count | Interpretation |
|---|---|---|---|
| Life science | explicitly-in-scope | Not reported | Three named internal directions are public, but tasks and counts are not. |
Relationship registry
Partial claims, non-evaluation uses, and third-party summaries stay visible without entering model comparisons.
evaluation
Benchmark: Anthropic Computational Biology Eval · version reported-2026-01-11
Work: Advancing Claude in healthcare and the life sciences · source version anthropic-healthcare-life-sciences-2026-01-11
Not reported / unresolved: task count; task data; grader; prompt; tools; repeats; absolute accuracy
Only the exact Opus 4.5 improvement annotation is normalized; plotted absolute values are not estimated.
evaluation
Benchmark: Anthropic Protein Understanding Eval · version reported-2026-01-11
Work: Advancing Claude in healthcare and the life sciences · source version anthropic-healthcare-life-sciences-2026-01-11
Not reported / unresolved: task count; task data; grader; prompt; tools; repeats; absolute accuracy
Only the exact Opus 4.5 improvement annotation is normalized; plotted absolute values are not estimated.
evaluation
Benchmark: Anthropic Scientific Figure Interpretation Eval · version reported-2026-01-11
Work: Advancing Claude in healthcare and the life sciences · source version anthropic-healthcare-life-sciences-2026-01-11
Not reported / unresolved: task count; task data; grader; prompt; tools; repeats; absolute accuracy
Only the exact Opus 4.5 improvement annotation is normalized; plotted absolute values are not estimated.
benchmark creation
Work: Advancing Claude in healthcare and the life sciences · source version anthropic-healthcare-life-sciences-2026-01-11
Anthropic is the creator and evaluator of this private internal suite; child-track evaluation relations are recorded separately.
Evaluation registry
A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.
Evaluation run
Evaluated models / systems: Claude Opus 4.1, Claude Opus 4.5, Claude Sonnet 4.5
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5 | delta baseline: Claude Opus 4.1 | percent delta as annotated | Not reported | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude Opus 4.5 | Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5 | Δ 10.5 percent delta as annotated Exact +10.5% chart annotation; no absolute accuracy or relative-versus-percentage-point interpretation is inferred. | Not reported |
Evaluation run
Evaluated models / systems: Claude Opus 4.1, Claude Opus 4.5, Claude Sonnet 4.5
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5 | delta baseline: Claude Opus 4.1 | percent delta as annotated | Not reported | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude Opus 4.5 | Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5 | Δ 10.3 percent delta as annotated Exact +10.3% chart annotation; no absolute accuracy or relative-versus-percentage-point interpretation is inferred. | Not reported |
Evaluation run
Evaluated models / systems: Claude Opus 4.1, Claude Opus 4.5, Claude Sonnet 4.5
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5 | delta baseline: Claude Opus 4.1 | percent delta as annotated | Not reported | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude Opus 4.5 | Accuracy improvement from Claude Opus 4.1 to Claude Opus 4.5 | Δ 13.2 percent delta as annotated Exact +13.2% chart annotation; no absolute accuracy or relative-versus-percentage-point interpretation is inferred. | Not reported |
Source locators remain visible; expand an item to inspect the exact Registry fields it supports.
/name/aliases/summary/kind/parent_id/organizations/release_date/latest_version/domains/capabilities/modalities/task_formats/task_counts/total/task_counts/basis/task_counts/subsets/coverage_notes/access/level/access/tasks/access/artifacts/access/grader/access/license/access/biosafety_notes/resources/implementations/versions/0/release_date/versions/0/as_of/versions/0/task_counts/total/versions/0/task_counts/basis/versions/0/task_counts/subsets/versions/0/formal_tracks