preprint · benchmark creator

Agentic systems are adept at solving well-scoped, verifiable problems in computational biology

Genentech · Roche · 2026-04-09

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

No BenchmarkUse relation is normalized for this legacy work yet. Existing EvaluationRuns remain available below.

Normalized evaluation runs

CompBioBench9 runs

Open benchmark record →

compbiobench-v1-hardest-codex-xhigh-three-runsvv1

Evaluated models / systems: Codex CLI (GPT-5.4)

Scopesubset · n=17
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / effortxhigh
Browserweb access through agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes initially; timeout rerun once with 240 minutes
TemperatureNot reported
SeedNot reported
Repeats3
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticsaccuracy averaged over three runs within the 17-task subset
Contaminationbenchmark construction reduces lookup shortcuts
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracy — difficulty Levels 4–5absolutepercentmean problem accuracy over three runswhitespace-stripped exact match

Results

ModelMetricValuen
Codex CLI (GPT-5.4)Accuracy — difficulty Levels 4–559 percent
Supplementary Figure 2 label; average across three runs.
17

Evidence

  • page: PDF pp. 13–14, Methods (Exact Codex configuration and common agent protocol.) — supports /protocol
  • figure: PDF p. 11, Supplementary Figure 2B and caption (Levels 4 and 5 grouped; 17 tasks; Codex label 59; three-run average.) — supports /scope, /metrics, /results
compbiobench-v1-codex-xhigh-three-runsvv1

Evaluated models / systems: Codex CLI (GPT-5.4)

Scopefull · n=100
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / effortxhigh
Browserweb access through the agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes initially; timed-out questions rerun clean once with 240 minutes
TemperatureNot reported
SeedNot reported
Repeats3
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticsmean over three independent runs; consistency reported as correct in all three and correct at least once
Contaminationsynthetic/augmented inputs and scrubbed identifiers reduce lookup shortcuts
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmean problem accuracy over three runswhitespace-stripped exact match
Wall-clock time per questionabsolutesecondsmean over questions and runs after 7200-second display clippingNot reported
Cost per questionabsoluteUSDmean over questions and runs after USD-10 display clippingNot reported

Results

ModelMetricValuen
Codex CLI (GPT-5.4)Accuracy83.3 percent
Mean of three full runs; 73% solved on all three and 92% at least once.
100
Codex CLI (GPT-5.4)Wall-clock time per question679 seconds
Printed Figure 2 label; mean across three runs.
100
Codex CLI (GPT-5.4)Cost per question1 USD
Printed Figure 2 label; mean across three runs.
100

Evidence

  • page: PDF pp. 13–14, Methods: agent execution and model configurations (Reports public wrapper prompt, internet/code/tool access, Conda isolation, 120/240-minute policy, Codex CLI v0.115.0, gpt-5.4, xhigh, and three runs.) — supports /benchmark_version, /scope, /protocol
  • figure: PDF pp. 3–4, Figure 2A–C and Results (Printed labels report 83.3% accuracy, 679.0 seconds, and USD 1.0; text reports three-run consistency.) — supports /metrics, /results
compbiobench-v1-haiku-one-run-120mvv1

Evaluated models / systems: Claude Code (Haiku 4.5)

Scopefull · n=100
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / effortunsupported
Browserweb access through the agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes; no 240-minute timeout rerun
TemperatureNot reported
SeedNot reported
Repeats1
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticssingle full-benchmark run
Contaminationsynthetic/augmented inputs and scrubbed identifiers reduce lookup shortcuts
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentproblem-weighted meanwhitespace-stripped exact match
Wall-clock time per questionabsolutesecondsmean after 7200-second display clippingNot reported
Cost per questionabsoluteUSDmean after USD-10 display clippingNot reported

Results

ModelMetricValuen
Claude Code (Haiku 4.5)Accuracy34 percent
One full run.
100
Claude Code (Haiku 4.5)Wall-clock time per question809.3 seconds
Printed Figure 2 label.
100
Claude Code (Haiku 4.5)Cost per question0.3 USD
Printed Figure 2 label.
100

Evidence

  • page: PDF pp. 13–14, Methods (Reports Claude Code v2.1.87, dated Haiku identifier, unsupported effort, one run, and no timeout rerun.) — supports /benchmark_version, /scope, /protocol
  • figure: PDF pp. 3–4, Figure 2A–C (Printed labels report 34.0%, 809.3 seconds, and USD 0.3.) — supports /metrics, /results
compbiobench-v1-hardest-haiku-one-run-120mvv1

Evaluated models / systems: Claude Code (Haiku 4.5)

Scopesubset · n=17
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / effortunsupported
Browserweb access through agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes; no 240-minute timeout rerun
TemperatureNot reported
SeedNot reported
Repeats1
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticsproblem-weighted accuracy in one run over the 17-task subset
Contaminationbenchmark construction reduces lookup shortcuts
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracy — difficulty Levels 4–5absolutepercentproblem-weighted meanwhitespace-stripped exact match

Results

ModelMetricValuen
Claude Code (Haiku 4.5)Accuracy — difficulty Levels 4–512 percent
Supplementary Figure 2 label; one run.
17

Evidence

  • page: PDF pp. 13–14, Methods (Exact Haiku configuration and timeout exception.) — supports /protocol
  • figure: PDF p. 11, Supplementary Figure 2B and caption (Levels 4 and 5 grouped; 17 tasks; Haiku label 12.) — supports /scope, /metrics, /results
compbiobench-v1-nonagentic-api-three-calls-no-filesvv1

Evaluated models / systems: ChatGPT 5.2, Claude Opus 4.6

Scopefull · n=100
Shotszero-shot
Turnssingle-turn API call
System prompt publicNot reported
Reasoning / effortdefault API parameters
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
Temperaturedefault
SeedNot reported
Repeats3
Graderwhitespace-stripped exact string match · human review: not reported
Statisticsmean accuracy over three calls per question
Contaminationbenchmark construction reduces lookup shortcuts
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmean over three calls per problemwhitespace-stripped exact match

Results

ModelMetricValuen
ChatGPT 5.2Accuracy5.3 percent
ChatGPT 5.2 non-agentic API baseline; three calls per question.
100
Claude Opus 4.6Accuracy3.7 percent
Claude Opus 4.6 non-agentic API baseline; three calls per question.
100

Evidence

  • page: PDF p. 13, Methods: LLM-only baselines (Default API parameters, three calls per question, no files, public prompt, and full benchmark.) — supports /benchmark_version, /scope, /protocol
  • figure: PDF pp. 3–4, Figure 2A and Results (Figure labels and text report ChatGPT 5.2 at 5.3% and Claude Opus 4.6 at 3.7%.) — supports /metrics, /results
compbiobench-v1-opus-max-three-runsvv1

Evaluated models / systems: Claude Code (Opus 4.6)

Scopefull · n=100
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / effortmax
Browserweb access through the agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes initially; timed-out questions rerun clean once with 240 minutes
TemperatureNot reported
SeedNot reported
Repeats3
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticsmean over three independent runs; consistency reported as correct in all three and at least once
Contaminationsynthetic/augmented inputs and scrubbed identifiers reduce lookup shortcuts
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmean problem accuracy over three runswhitespace-stripped exact match
Wall-clock time per questionabsolutesecondsmean after 7200-second display clippingNot reported
Cost per questionabsoluteUSDmean after USD-10 display clippingNot reported

Results

ModelMetricValuen
Claude Code (Opus 4.6)Accuracy81 percent
Mean of three full runs; 73% solved all three and 86% at least once.
100
Claude Code (Opus 4.6)Wall-clock time per question1101 seconds
Printed Figure 2 label.
100
Claude Code (Opus 4.6)Cost per question1.7 USD
Printed Figure 2 label.
100

Evidence

  • page: PDF pp. 13–14, Methods: agent execution and model configurations (Reports Claude Code v2.1.87, claude-opus-4-6 1M, max effort, three runs, tools, and timeout policy.) — supports /benchmark_version, /scope, /protocol
  • figure: PDF pp. 3–4, Figure 2A–C and Results (Reports 81.0% accuracy, 1101.0 seconds, USD 1.7, and consistency.) — supports /metrics, /results
compbiobench-v1-hardest-opus-max-three-runsvv1

Evaluated models / systems: Claude Code (Opus 4.6)

Scopesubset · n=17
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / effortmax
Browserweb access through agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes initially; timeout rerun once with 240 minutes
TemperatureNot reported
SeedNot reported
Repeats3
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticsaccuracy averaged over three runs within the 17-task subset
Contaminationbenchmark construction reduces lookup shortcuts
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracy — difficulty Levels 4–5absolutepercentmean problem accuracy over three runswhitespace-stripped exact match

Results

ModelMetricValuen
Claude Code (Opus 4.6)Accuracy — difficulty Levels 4–569 percent
Supplementary Figure 2 label; average across three runs.
17

Evidence

  • page: PDF pp. 13–14, Methods (Exact Claude Code Opus configuration and common protocol.) — supports /protocol
  • figure: PDF p. 11, Supplementary Figure 2B and caption (Levels 4 and 5 grouped; 17 tasks; Opus label 69; three-run average.) — supports /scope, /metrics, /results
compbiobench-v1-sonnet-high-one-runvv1

Evaluated models / systems: Claude Code (Sonnet 4.6)

Scopefull · n=100
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / efforthigh
Browserweb access through the agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes initially; timed-out questions rerun clean once with 240 minutes
TemperatureNot reported
SeedNot reported
Repeats1
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticssingle full-benchmark run
Contaminationsynthetic/augmented inputs and scrubbed identifiers reduce lookup shortcuts
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentproblem-weighted meanwhitespace-stripped exact match
Wall-clock time per questionabsolutesecondsmean after 7200-second display clippingNot reported
Cost per questionabsoluteUSDmean after USD-10 display clippingNot reported

Results

ModelMetricValuen
Claude Code (Sonnet 4.6)Accuracy70 percent
One full run.
100
Claude Code (Sonnet 4.6)Wall-clock time per question1049.2 seconds
Printed Figure 2 label.
100
Claude Code (Sonnet 4.6)Cost per question1.2 USD
Printed Figure 2 label.
100

Evidence

  • page: PDF pp. 13–14, Methods (Reports Claude Code v2.1.87, claude-sonnet-4-6 1M, high effort, one run, tools, and timeout policy.) — supports /benchmark_version, /scope, /protocol
  • figure: PDF pp. 3–4, Figure 2A–C (Printed labels report 70.0%, 1049.2 seconds, and USD 1.2.) — supports /metrics, /results
compbiobench-v1-hardest-sonnet-high-one-runvv1

Evaluated models / systems: Claude Code (Sonnet 4.6)

Scopesubset · n=17
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / efforthigh
Browserweb access through agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes initially; timeout rerun once with 240 minutes
TemperatureNot reported
SeedNot reported
Repeats1
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticsproblem-weighted accuracy in one run over the 17-task subset
Contaminationbenchmark construction reduces lookup shortcuts
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracy — difficulty Levels 4–5absolutepercentproblem-weighted meanwhitespace-stripped exact match

Results

ModelMetricValuen
Claude Code (Sonnet 4.6)Accuracy — difficulty Levels 4–553 percent
Supplementary Figure 2 label; one run.
17

Evidence

  • page: PDF pp. 13–14, Methods (Exact Claude Code Sonnet configuration and common protocol.) — supports /protocol
  • figure: PDF p. 11, Supplementary Figure 2B and caption (Levels 4 and 5 grouped; 17 tasks; Sonnet label 53.) — supports /scope, /metrics, /results