agentic-eval · audited · verified 2026-07-21

BioMysteryBench

An agentic bioinformatics benchmark of objective, expert-authored mysteries over anonymized real-world biological data, scored on final answers rather than prescribed analysis paths.

+6 more
Version and evaluation audit: the current v11 release has 90 problems (73 human-solvable, 17 human-hard) after nine removals and 24 problem edits. The April model results remain tied to v8's 99 problems (76/23), with five episodes per model–problem pair; they are never presented as v11 scores.

Benchmark definition

What is counted

Version
v11
Total
90 (v11 mystery-bioinformatics problems after the June 2026 answer-key audit)
Task formats
open-ended bioinformatics investigation; containerized agent episode
Capabilities
Data analysisCodingTool useScientific reasoning
Modalities
TextTableDNA or RNA sequence3D structureRaw omicsDatabaseCode

Version history

VersionStatusRelease / as-ofTotalFormal tracks
v8
biomysterybench-v8
superseded2026-04-2899 (initial public-release problems after pre-release QC exclusions)None registered
v11
biomysterybench-v11
current2026-07-0690 (v11 mystery-bioinformatics problems after the June 2026 answer-key audit)None registered

Tracks and subsets

IDCountBasisPartition?Notes
Human-solvable (v11)
human-solvable
73problems solved by at least one human benchmarkerExclusive & exhaustiveThe v11 changelog reports 73 human-solvable problems.
Human-hard (v11)
human-difficult
17problems not solved by the human panelExclusive & exhaustiveThe v11 changelog calls this split human-hard; the April evaluation report used human-difficult for its 23-problem predecessor.

Scientific Task Atlas

Scientific task classification

partial for v11. Official modality examples support a broad end-to-end analysis mapping but not exhaustive task-topic counts.

Scientific taskCoverageCountMappingEvidence
End-to-end computational analysisexplicitly-in-scope90 problems
v11 mystery-bioinformatics problems after the June 2026 answer-key audit
official-taxonomy
high confidence
biomysterybench-evidence-v11
Each mystery is scored on its final answer rather than a prescribed analysis path.
Omics and cellular analysisexplicitly-in-scopeNot reported
v11 mystery-bioinformatics problems using omics and cellular data.
official-taxonomy
high confidence
biomysterybench-evidence-taxonomy
Single-cell, proteomics, metabolomics, and other modalities are explicit, without a leaf-task count.

Scientific coverage notes

DomainCoverageCountInterpretation
Protein structureexplicitly-in-scopeNot reportedThe official report gives crystal-structure organism identification as an example, but no standalone structure-task count.
Single-cellexplicitly-in-scopeNot reportedThe official report explicitly lists scRNA-seq and gives a single-cell organ-identification example; no standalone count is published.
Proteomicsexplicitly-in-scopeNot reportedThe official report says several questions use proteomics, but does not publish a standalone count.
Metabolomicsexplicitly-in-scopeNot reportedThe official report says several questions use metabolomics, but does not publish a standalone count.

Relationship registry

How works use this benchmark

Partial claims, non-evaluation uses, and third-party summaries stay visible without entering model comparisons.

Partial evaluation claims

evaluation

system-card-claude-opus-5-biomysterybench-1-use

Partialsubset

Work: System Card: Claude Opus 5 · source version system-card-claude-opus-5-2026-07-24

Selection
formal subset · Human Difficult: problems unsolved by humans with an objective ground-truth solution
Metrics
Score
Linked runs
None

Not reported / unresolved: Exact benchmark version is not reported.; Overall total and current Human Solvable and Human Difficult subset sizes are not reported; only removal counts are given.; Prompt, shots, reasoning settings, budget, seed, repeats, grader, and human review are not reported.; Score definition and aggregation are not reported.; benchmark version; realized n/scope

AI-assisted double-pass extraction; values are limited to independently supported claims.

Evidence
  • section: Section 8.17.1 BioMysteryBench
    Supports: /relation_type
  • section: Section 8.17.1
    Supports: /benchmark_id
  • section: Section 8.17.1 BioMysteryBench
    Supports: /scope
  • section: Section 8.17.1 BioMysteryBench
    Supports: /scope
  • section: Section 8.17.1 BioMysteryBench
    Supports: /scope
  • section: Section 8.17.1 BioMysteryBench
    Supports: /scope
  • section: Section 8.17.1 BioMysteryBench
    Supports: /scope
  • figure: Figure 8.17.6.A, BioMysteryBench panel
    Supports: /metric_labels
  • section: Section 8.17.1 BioMysteryBench
    Supports: /model_ids
  • section: Section 8.17.1 BioMysteryBench
    Supports: /model_ids
  • section: Section 8.17.1 BioMysteryBench
    Supports: /model_ids
  • section: Section 8.17.1 BioMysteryBench
    Supports: /model_ids

Evaluation registry

Works and run settings

A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.

biomysterybench-v8-full-five-episodes-protocolvv8

Evaluated models / systems: Claude Haiku 4.5, Claude Mythos Preview, Claude Opus 4.6, Claude Opus 4.7, Claude Sonnet 4.6

Scopefull · n=99
ShotsNot reported
Turnsmulti-turn agent episode
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNot reported
Internetallowlisted external access
DatabasesNCBI, Ensembl, other canonical bioinformatics databases allowed per problem
Code executionYes
ContainerYes
External toolspreinstalled canonical bioinformatics tools plus packages installable through pip and conda
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats5
Graderobjective final-answer scoring against an expert-authored answer rubric · human review: not reported
Statisticsaccuracy averaged over five episodes per problem with error bars from bootstrap sampling within problems; per-problem reliability analyzed by 0-of-5 through 5-of-5 solve counts
Contaminationsource-dataset reverse identification prohibited; accession-ID leaks scrubbed from 16 v8 problems
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmean over five trials per problem, then mean over the selected problemsexpert-authored answer rubric

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • section: “Benchmarking models on verifiable biological tasks,” “Human-solvable,” and “Human-difficult” (Reports 99 problems partitioned into 76 and 23 after four failed-QC candidates were removed.) — supports /benchmark_version, /scope
  • section: Environment description, method-agnostic property, Figures 1–3 captions, and continuing reliability analysis (Reports containers, databases, package installation, final-answer grading, five episodes, bootstrap-within-problem error bars, and solve-count reliability profiles.) — supports /protocol/turns, /protocol/tools/internet, /protocol/tools/databases, /protocol/tools/code_execution, /protocol/tools/container, /protocol/tools/external_tools, /protocol/repeats, /protocol/grader, /protocol/statistical, /metrics
  • repository-path: README.md Rules and CHANGELOG.md v8 entry at commit 51c9024021b8989a0cb06ae623b02f90d14c2da3 (Documents prohibited accession/reverse lookup, permitted standard database use, and 16 scrubbed v8 problems.) — supports /protocol/contamination
  • section: Complete public report and figure captions (The public protocol does not report shots, a browser, a system prompt, effort, budgets, temperature, seed, human-review status, exact grader implementation, bootstrap count, or numeric confidence bounds.) — supports /protocol/shots, /protocol/system_prompt_public, /protocol/reasoning, /protocol/tools/browser, /protocol/token_budget, /protocol/time_budget, /protocol/temperature, /protocol/seed

Evaluation run

biomysterybench-v8-human-difficult

From Evaluating Claude's bioinformatics research capabilities with BioMysteryBench

biomysterybench-v8-human-difficult-five-episodesvv8

Evaluated models / systems: Claude Haiku 4.5, Claude Mythos Preview, Claude Opus 4.6, Claude Opus 4.7, Claude Sonnet 4.6

Scopesubset · n=23
ShotsNot reported
Turnsmulti-turn agent episode
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNot reported
Internetallowlisted external access
Databasescanonical bioinformatics databases including NCBI and Ensembl
Code executionYes
ContainerYes
External toolspreinstalled tools plus packages installable through pip and conda
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats5
Graderobjective final-answer scoring against an expert-authored answer rubric · human review: not reported
Statisticsaccuracy averaged over five episodes per problem with error bars from bootstrap sampling within problems; per-problem reliability also reported by 0-of-5 through 5-of-5 solve counts
Contaminationsource-dataset reverse identification prohibited; accession-ID leaks scrubbed from 16 v8 problems
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracy — human-difficultabsolutepercentmean over five episodes per problem, then mean over 23 problemsexpert-authored answer rubric

Results

ModelMetricValuen
Claude Haiku 4.5Accuracy — human-difficult5.2 percent
115 model episodes; official figure labels the point estimate and draws unlabeled bootstrap error bars.
23
Claude Sonnet 4.6Accuracy — human-difficult19.1 percent
115 model episodes; official figure labels the point estimate and draws unlabeled bootstrap error bars.
23
Claude Opus 4.6Accuracy — human-difficult23.5 percent
115 model episodes; official figure labels the point estimate and draws unlabeled bootstrap error bars.
23
Claude Opus 4.7Accuracy — human-difficult27 percent
115 model episodes; official figure labels the point estimate and draws unlabeled bootstrap error bars.
23
Claude Mythos PreviewAccuracy — human-difficult29.6 percent
115 model episodes; official figure labels the point estimate and draws unlabeled bootstrap error bars.
23

Evidence

  • section: “Human-difficult” and Figure 2 caption (Defines the 23-problem subset, five episodes, and within-problem bootstrap.) — supports /scope, /protocol/repeats, /protocol/statistical, /metrics
  • figure: Figure 2: BioMysteryBench human-difficult set (23 problems) (Printed labels report 5.2, 19.1, 23.5, 27.0, and 29.6 percent.) — supports /results
  • section: Environment description and method-agnostic property (Reports the agent trajectory, container, databases, code/tool access, package installation, and final-answer rubric grading.) — supports /protocol/turns, /protocol/tools/internet, /protocol/tools/databases, /protocol/tools/code_execution, /protocol/tools/container, /protocol/tools/external_tools, /protocol/grader
  • repository-path: README.md Rules and CHANGELOG.md v8 entry at commit 51c9024021b8989a0cb06ae623b02f90d14c2da3 (Documents prohibited reverse lookup, permitted standard database use, and 16 scrubbed v8 problems.) — supports /protocol/contamination
  • section: Complete public report and Figure 2 caption (Does not report shots, browser interface, system prompt, effort, budgets, temperature, seed, human-review status, bootstrap count, or numeric confidence bounds.) — supports /protocol/shots, /protocol/system_prompt_public, /protocol/reasoning, /protocol/tools/browser, /protocol/token_budget, /protocol/time_budget, /protocol/temperature, /protocol/seed

Evaluation run

biomysterybench-v8-human-solvable

From Evaluating Claude's bioinformatics research capabilities with BioMysteryBench

biomysterybench-v8-human-solvable-five-episodesvv8

Evaluated models / systems: Claude Haiku 4.5, Claude Mythos Preview, Claude Opus 4.6, Claude Opus 4.7, Claude Sonnet 4.6

Scopesubset · n=76
ShotsNot reported
Turnsmulti-turn agent episode
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNot reported
Internetallowlisted external access
Databasescanonical bioinformatics databases including NCBI and Ensembl
Code executionYes
ContainerYes
External toolspreinstalled tools plus packages installable through pip and conda
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats5
Graderobjective final-answer scoring against an expert-authored answer rubric · human review: not reported
Statisticsaccuracy averaged over five trials per problem with error bars from bootstrap sampling within problems; per-problem reliability also reported by 0-of-5 through 5-of-5 solve counts
Contaminationsource-dataset reverse identification prohibited; accession-ID leaks scrubbed from 16 v8 problems
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracy — human-solvableabsolutepercentmean over five trials per problem, then mean over 76 problemsexpert-authored answer rubric

Results

ModelMetricValuen
Claude Haiku 4.5Accuracy — human-solvable36.8 percent
380 model episodes; official figure labels the point estimate and draws unlabeled bootstrap error bars.
76
Claude Sonnet 4.6Accuracy — human-solvable71.8 percent
380 model episodes; official figure labels the point estimate and draws unlabeled bootstrap error bars.
76
Claude Opus 4.6Accuracy — human-solvable77.4 percent
380 model episodes; official figure labels the point estimate and draws unlabeled bootstrap error bars.
76
Claude Opus 4.7Accuracy — human-solvable78.9 percent
380 model episodes; official figure labels the point estimate and draws unlabeled bootstrap error bars.
76
Claude Mythos PreviewAccuracy — human-solvable82.6 percent
380 model episodes; official figure labels the point estimate and draws unlabeled bootstrap error bars.
76

Evidence

  • section: “Human-solvable” and Figure 1 caption (Defines the 76-problem subset, five trials, and within-problem bootstrap.) — supports /scope, /protocol/repeats, /protocol/statistical, /metrics
  • figure: Figure 1: BioMysteryBench human-solvable set (76 problems) (Printed labels report 36.8, 71.8, 77.4, 78.9, and 82.6 percent.) — supports /results
  • section: Environment description and method-agnostic property (Reports the agent trajectory, container, databases, code/tool access, package installation, and final-answer rubric grading.) — supports /protocol/turns, /protocol/tools/internet, /protocol/tools/databases, /protocol/tools/code_execution, /protocol/tools/container, /protocol/tools/external_tools, /protocol/grader
  • repository-path: README.md Rules and CHANGELOG.md v8 entry at commit 51c9024021b8989a0cb06ae623b02f90d14c2da3 (Documents prohibited reverse lookup, permitted standard database use, and 16 scrubbed v8 problems.) — supports /protocol/contamination
  • section: Complete public report and Figure 1 caption (Does not report shots, browser interface, system prompt, effort, budgets, temperature, seed, human-review status, bootstrap count, or numeric confidence bounds.) — supports /protocol/shots, /protocol/system_prompt_public, /protocol/reasoning, /protocol/tools/browser, /protocol/token_budget, /protocol/time_budget, /protocol/temperature, /protocol/seed

Comparable result views

Accuracy — human-difficult

biomysterybench-v8-human-difficult · biomysterybench-v8-human-difficult-five-episodes

CSV ↓
Accessible data table
ModelValueComparability group
Claude Haiku 4.55.2biomysterybench-v8-human-difficult-five-episodes
Claude Sonnet 4.619.1biomysterybench-v8-human-difficult-five-episodes
Claude Opus 4.623.5biomysterybench-v8-human-difficult-five-episodes
Claude Opus 4.727biomysterybench-v8-human-difficult-five-episodes
Claude Mythos Preview29.6biomysterybench-v8-human-difficult-five-episodes

Accuracy — human-solvable

biomysterybench-v8-human-solvable · biomysterybench-v8-human-solvable-five-episodes

CSV ↓
Accessible data table
ModelValueComparability group
Claude Haiku 4.536.8biomysterybench-v8-human-solvable-five-episodes
Claude Sonnet 4.671.8biomysterybench-v8-human-solvable-five-episodes
Claude Opus 4.677.4biomysterybench-v8-human-solvable-five-episodes
Claude Opus 4.778.9biomysterybench-v8-human-solvable-five-episodes
Claude Mythos Preview82.6biomysterybench-v8-human-solvable-five-episodes

Evidence and change history

Source locators remain visible; expand an item to inspect the exact Registry fields it supports.

Evaluating Claude's bioinformatics research capabilities with BioMysteryBench · section: Title, publication date, and “Benchmarking models on verifiable biological tasks with BioMysteryBench” (Identifies Anthropic as creator and defines the method-agnostic, objective-answer agent benchmark.) · Supports 7 fields

Open source →

  • /name
  • /organizations
  • /release_date
  • /kind
  • /summary
  • /capabilities
  • /task_formats
Evaluating Claude's bioinformatics research capabilities with BioMysteryBench · section: “Benchmarking models on verifiable biological tasks,” “Human-solvable,” and “Human-difficult”; v8 entry in the official dataset changelog (Reports the original 99 total, 76 human-solvable, 23 human-difficult, and four pre-release QC exclusions.) · Supports 3 fields

Open source →

  • /versions/0/task_counts/total
  • /versions/0/task_counts/basis
  • /versions/0/task_counts/subsets
biomysterybench-preview-resource · repository-path: CHANGELOG.md, v11 (2026-07-06), at commit 51c9024021b8989a0cb06ae623b02f90d14c2da3 (Reports 99 to 90, the 73/17 partition, nine removals, 24 modified problems, and revised all-or-nothing rubric language.) · Supports 8 fields

Open source →

  • /latest_version
  • /task_counts/total
  • /task_counts/basis
  • /task_counts/subsets
  • /versions/1/task_counts/total
  • /versions/1/task_counts/basis
  • /versions/1/task_counts/subsets
  • /scientific_task_classification/entries/0
Evaluating Claude's bioinformatics research capabilities with BioMysteryBench · section: “Example questions” and the environment/property list (Explicitly lists WGS, RNA-seq, scRNA-seq, methylation, ChIP-seq, metagenomics, Hi-C, proteomics, metabolomics, crystal structures, databases, and coding/tool use.) · Supports 7 fields

Open source →

  • /domains
  • /modalities
  • /coverage_notes/0
  • /coverage_notes/1
  • /coverage_notes/2
  • /coverage_notes/3
  • /scientific_task_classification/entries/1
biomysterybench-full-resource · dataset-card: Access conditions, Contents, Rules, and License and terms of use at commit b5a889c4757214ec9a6ade876b734f920a7799db (Documents the gated 90-problem release, evaluation-only restriction, CC BY 4.0 benchmark materials, source-policy data archives, and answer-rubric fields.) · Supports 3 fields

Open source →

  • /access/level
  • /access/license
  • /resources

View source-level modification history on GitHub →