official-release · benchmark creator

Evaluating Claude's bioinformatics research capabilities with BioMysteryBench

Anthropic · 2026-04-29

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

No BenchmarkUse relation is normalized for this legacy work yet. Existing EvaluationRuns remain available below.

Normalized evaluation runs

BioMysteryBench3 runs

Open benchmark record →

biomysterybench-v8-full-five-episodes-protocolvv8

Evaluated models / systems: Claude Haiku 4.5, Claude Mythos Preview, Claude Opus 4.6, Claude Opus 4.7, Claude Sonnet 4.6

Scopefull · n=99
ShotsNot reported
Turnsmulti-turn agent episode
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNot reported
Internetallowlisted external access
DatabasesNCBI, Ensembl, other canonical bioinformatics databases allowed per problem
Code executionYes
ContainerYes
External toolspreinstalled canonical bioinformatics tools plus packages installable through pip and conda
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats5
Graderobjective final-answer scoring against an expert-authored answer rubric · human review: not reported
Statisticsaccuracy averaged over five episodes per problem with error bars from bootstrap sampling within problems; per-problem reliability analyzed by 0-of-5 through 5-of-5 solve counts
Contaminationsource-dataset reverse identification prohibited; accession-ID leaks scrubbed from 16 v8 problems
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmean over five trials per problem, then mean over the selected problemsexpert-authored answer rubric

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • section: “Benchmarking models on verifiable biological tasks,” “Human-solvable,” and “Human-difficult” (Reports 99 problems partitioned into 76 and 23 after four failed-QC candidates were removed.) — supports /benchmark_version, /scope
  • section: Environment description, method-agnostic property, Figures 1–3 captions, and continuing reliability analysis (Reports containers, databases, package installation, final-answer grading, five episodes, bootstrap-within-problem error bars, and solve-count reliability profiles.) — supports /protocol/turns, /protocol/tools/internet, /protocol/tools/databases, /protocol/tools/code_execution, /protocol/tools/container, /protocol/tools/external_tools, /protocol/repeats, /protocol/grader, /protocol/statistical, /metrics
  • repository-path: README.md Rules and CHANGELOG.md v8 entry at commit 51c9024021b8989a0cb06ae623b02f90d14c2da3 (Documents prohibited accession/reverse lookup, permitted standard database use, and 16 scrubbed v8 problems.) — supports /protocol/contamination
  • section: Complete public report and figure captions (The public protocol does not report shots, a browser, a system prompt, effort, budgets, temperature, seed, human-review status, exact grader implementation, bootstrap count, or numeric confidence bounds.) — supports /protocol/shots, /protocol/system_prompt_public, /protocol/reasoning, /protocol/tools/browser, /protocol/token_budget, /protocol/time_budget, /protocol/temperature, /protocol/seed

Evaluation run

biomysterybench-v8-human-difficult

From Evaluating Claude's bioinformatics research capabilities with BioMysteryBench

biomysterybench-v8-human-difficult-five-episodesvv8

Evaluated models / systems: Claude Haiku 4.5, Claude Mythos Preview, Claude Opus 4.6, Claude Opus 4.7, Claude Sonnet 4.6

Scopesubset · n=23
ShotsNot reported
Turnsmulti-turn agent episode
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNot reported
Internetallowlisted external access
Databasescanonical bioinformatics databases including NCBI and Ensembl
Code executionYes
ContainerYes
External toolspreinstalled tools plus packages installable through pip and conda
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats5
Graderobjective final-answer scoring against an expert-authored answer rubric · human review: not reported
Statisticsaccuracy averaged over five episodes per problem with error bars from bootstrap sampling within problems; per-problem reliability also reported by 0-of-5 through 5-of-5 solve counts
Contaminationsource-dataset reverse identification prohibited; accession-ID leaks scrubbed from 16 v8 problems
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracy — human-difficultabsolutepercentmean over five episodes per problem, then mean over 23 problemsexpert-authored answer rubric

Results

ModelMetricValuen
Claude Haiku 4.5Accuracy — human-difficult5.2 percent
115 model episodes; official figure labels the point estimate and draws unlabeled bootstrap error bars.
23
Claude Sonnet 4.6Accuracy — human-difficult19.1 percent
115 model episodes; official figure labels the point estimate and draws unlabeled bootstrap error bars.
23
Claude Opus 4.6Accuracy — human-difficult23.5 percent
115 model episodes; official figure labels the point estimate and draws unlabeled bootstrap error bars.
23
Claude Opus 4.7Accuracy — human-difficult27 percent
115 model episodes; official figure labels the point estimate and draws unlabeled bootstrap error bars.
23
Claude Mythos PreviewAccuracy — human-difficult29.6 percent
115 model episodes; official figure labels the point estimate and draws unlabeled bootstrap error bars.
23

Evidence

  • section: “Human-difficult” and Figure 2 caption (Defines the 23-problem subset, five episodes, and within-problem bootstrap.) — supports /scope, /protocol/repeats, /protocol/statistical, /metrics
  • figure: Figure 2: BioMysteryBench human-difficult set (23 problems) (Printed labels report 5.2, 19.1, 23.5, 27.0, and 29.6 percent.) — supports /results
  • section: Environment description and method-agnostic property (Reports the agent trajectory, container, databases, code/tool access, package installation, and final-answer rubric grading.) — supports /protocol/turns, /protocol/tools/internet, /protocol/tools/databases, /protocol/tools/code_execution, /protocol/tools/container, /protocol/tools/external_tools, /protocol/grader
  • repository-path: README.md Rules and CHANGELOG.md v8 entry at commit 51c9024021b8989a0cb06ae623b02f90d14c2da3 (Documents prohibited reverse lookup, permitted standard database use, and 16 scrubbed v8 problems.) — supports /protocol/contamination
  • section: Complete public report and Figure 2 caption (Does not report shots, browser interface, system prompt, effort, budgets, temperature, seed, human-review status, bootstrap count, or numeric confidence bounds.) — supports /protocol/shots, /protocol/system_prompt_public, /protocol/reasoning, /protocol/tools/browser, /protocol/token_budget, /protocol/time_budget, /protocol/temperature, /protocol/seed

Evaluation run

biomysterybench-v8-human-solvable

From Evaluating Claude's bioinformatics research capabilities with BioMysteryBench

biomysterybench-v8-human-solvable-five-episodesvv8

Evaluated models / systems: Claude Haiku 4.5, Claude Mythos Preview, Claude Opus 4.6, Claude Opus 4.7, Claude Sonnet 4.6

Scopesubset · n=76
ShotsNot reported
Turnsmulti-turn agent episode
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNot reported
Internetallowlisted external access
Databasescanonical bioinformatics databases including NCBI and Ensembl
Code executionYes
ContainerYes
External toolspreinstalled tools plus packages installable through pip and conda
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats5
Graderobjective final-answer scoring against an expert-authored answer rubric · human review: not reported
Statisticsaccuracy averaged over five trials per problem with error bars from bootstrap sampling within problems; per-problem reliability also reported by 0-of-5 through 5-of-5 solve counts
Contaminationsource-dataset reverse identification prohibited; accession-ID leaks scrubbed from 16 v8 problems
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracy — human-solvableabsolutepercentmean over five trials per problem, then mean over 76 problemsexpert-authored answer rubric

Results

ModelMetricValuen
Claude Haiku 4.5Accuracy — human-solvable36.8 percent
380 model episodes; official figure labels the point estimate and draws unlabeled bootstrap error bars.
76
Claude Sonnet 4.6Accuracy — human-solvable71.8 percent
380 model episodes; official figure labels the point estimate and draws unlabeled bootstrap error bars.
76
Claude Opus 4.6Accuracy — human-solvable77.4 percent
380 model episodes; official figure labels the point estimate and draws unlabeled bootstrap error bars.
76
Claude Opus 4.7Accuracy — human-solvable78.9 percent
380 model episodes; official figure labels the point estimate and draws unlabeled bootstrap error bars.
76
Claude Mythos PreviewAccuracy — human-solvable82.6 percent
380 model episodes; official figure labels the point estimate and draws unlabeled bootstrap error bars.
76

Evidence

  • section: “Human-solvable” and Figure 1 caption (Defines the 76-problem subset, five trials, and within-problem bootstrap.) — supports /scope, /protocol/repeats, /protocol/statistical, /metrics
  • figure: Figure 1: BioMysteryBench human-solvable set (76 problems) (Printed labels report 36.8, 71.8, 77.4, 78.9, and 82.6 percent.) — supports /results
  • section: Environment description and method-agnostic property (Reports the agent trajectory, container, databases, code/tool access, package installation, and final-answer rubric grading.) — supports /protocol/turns, /protocol/tools/internet, /protocol/tools/databases, /protocol/tools/code_execution, /protocol/tools/container, /protocol/tools/external_tools, /protocol/grader
  • repository-path: README.md Rules and CHANGELOG.md v8 entry at commit 51c9024021b8989a0cb06ae623b02f90d14c2da3 (Documents prohibited reverse lookup, permitted standard database use, and 16 scrubbed v8 problems.) — supports /protocol/contamination
  • section: Complete public report and Figure 1 caption (Does not report shots, browser interface, system prompt, effort, budgets, temperature, seed, human-review status, bootstrap count, or numeric confidence bounds.) — supports /protocol/shots, /protocol/system_prompt_public, /protocol/reasoning, /protocol/tools/browser, /protocol/token_budget, /protocol/time_budget, /protocol/temperature, /protocol/seed