preprint · benchmark creator
Agentic systems are adept at solving well-scoped, verifiable problems in computational biology
Genentech · Roche · 2026-04-09
Relationship layer
Benchmark usage
This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.
No BenchmarkUse relation is normalized for this legacy work yet. Existing EvaluationRuns remain available below.
Normalized evaluation runs
CompBioBench9 runs
Open benchmark record →
compbiobench-v1-hardest-codex-xhigh-three-runsvv1
Evaluated models / systems: Codex CLI (GPT-5.4)
Scopesubset · n=17
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / effortxhigh
Browserweb access through agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes initially; timeout rerun once with 240 minutes
TemperatureNot reported
SeedNot reported
Repeats3
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticsaccuracy averaged over three runs within the 17-task subset
Contaminationbenchmark construction reduces lookup shortcuts
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Accuracy — difficulty Levels 4–5 | absolute | percent | mean problem accuracy over three runs | whitespace-stripped exact match |
Results
| Model | Metric | Value | n |
|---|
| Codex CLI (GPT-5.4) | Accuracy — difficulty Levels 4–5 | 59 percent Supplementary Figure 2 label; average across three runs. | 17 |
Evidence
- page: PDF pp. 13–14, Methods (Exact Codex configuration and common agent protocol.) — supports /protocol
- figure: PDF p. 11, Supplementary Figure 2B and caption (Levels 4 and 5 grouped; 17 tasks; Codex label 59; three-run average.) — supports /scope, /metrics, /results
compbiobench-v1-codex-xhigh-three-runsvv1
Evaluated models / systems: Codex CLI (GPT-5.4)
Scopefull · n=100
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / effortxhigh
Browserweb access through the agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes initially; timed-out questions rerun clean once with 240 minutes
TemperatureNot reported
SeedNot reported
Repeats3
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticsmean over three independent runs; consistency reported as correct in all three and correct at least once
Contaminationsynthetic/augmented inputs and scrubbed identifiers reduce lookup shortcuts
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Accuracy | absolute | percent | mean problem accuracy over three runs | whitespace-stripped exact match |
| Wall-clock time per question | absolute | seconds | mean over questions and runs after 7200-second display clipping | Not reported |
| Cost per question | absolute | USD | mean over questions and runs after USD-10 display clipping | Not reported |
Results
| Model | Metric | Value | n |
|---|
| Codex CLI (GPT-5.4) | Accuracy | 83.3 percent Mean of three full runs; 73% solved on all three and 92% at least once. | 100 |
| Codex CLI (GPT-5.4) | Wall-clock time per question | 679 seconds Printed Figure 2 label; mean across three runs. | 100 |
| Codex CLI (GPT-5.4) | Cost per question | 1 USD Printed Figure 2 label; mean across three runs. | 100 |
Evidence
- page: PDF pp. 13–14, Methods: agent execution and model configurations (Reports public wrapper prompt, internet/code/tool access, Conda isolation, 120/240-minute policy, Codex CLI v0.115.0, gpt-5.4, xhigh, and three runs.) — supports /benchmark_version, /scope, /protocol
- figure: PDF pp. 3–4, Figure 2A–C and Results (Printed labels report 83.3% accuracy, 679.0 seconds, and USD 1.0; text reports three-run consistency.) — supports /metrics, /results
compbiobench-v1-haiku-one-run-120mvv1
Evaluated models / systems: Claude Code (Haiku 4.5)
Scopefull · n=100
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / effortunsupported
Browserweb access through the agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes; no 240-minute timeout rerun
TemperatureNot reported
SeedNot reported
Repeats1
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticssingle full-benchmark run
Contaminationsynthetic/augmented inputs and scrubbed identifiers reduce lookup shortcuts
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Accuracy | absolute | percent | problem-weighted mean | whitespace-stripped exact match |
| Wall-clock time per question | absolute | seconds | mean after 7200-second display clipping | Not reported |
| Cost per question | absolute | USD | mean after USD-10 display clipping | Not reported |
Results
Evidence
- page: PDF pp. 13–14, Methods (Reports Claude Code v2.1.87, dated Haiku identifier, unsupported effort, one run, and no timeout rerun.) — supports /benchmark_version, /scope, /protocol
- figure: PDF pp. 3–4, Figure 2A–C (Printed labels report 34.0%, 809.3 seconds, and USD 0.3.) — supports /metrics, /results
compbiobench-v1-hardest-haiku-one-run-120mvv1
Evaluated models / systems: Claude Code (Haiku 4.5)
Scopesubset · n=17
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / effortunsupported
Browserweb access through agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes; no 240-minute timeout rerun
TemperatureNot reported
SeedNot reported
Repeats1
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticsproblem-weighted accuracy in one run over the 17-task subset
Contaminationbenchmark construction reduces lookup shortcuts
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Accuracy — difficulty Levels 4–5 | absolute | percent | problem-weighted mean | whitespace-stripped exact match |
Results
| Model | Metric | Value | n |
|---|
| Claude Code (Haiku 4.5) | Accuracy — difficulty Levels 4–5 | 12 percent Supplementary Figure 2 label; one run. | 17 |
Evidence
- page: PDF pp. 13–14, Methods (Exact Haiku configuration and timeout exception.) — supports /protocol
- figure: PDF p. 11, Supplementary Figure 2B and caption (Levels 4 and 5 grouped; 17 tasks; Haiku label 12.) — supports /scope, /metrics, /results
compbiobench-v1-nonagentic-api-three-calls-no-filesvv1
Evaluated models / systems: ChatGPT 5.2, Claude Opus 4.6
Scopefull · n=100
Shotszero-shot
Turnssingle-turn API call
System prompt publicNot reported
Reasoning / effortdefault API parameters
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
Temperaturedefault
SeedNot reported
Repeats3
Graderwhitespace-stripped exact string match · human review: not reported
Statisticsmean accuracy over three calls per question
Contaminationbenchmark construction reduces lookup shortcuts
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Accuracy | absolute | percent | mean over three calls per problem | whitespace-stripped exact match |
Results
| Model | Metric | Value | n |
|---|
| ChatGPT 5.2 | Accuracy | 5.3 percent ChatGPT 5.2 non-agentic API baseline; three calls per question. | 100 |
| Claude Opus 4.6 | Accuracy | 3.7 percent Claude Opus 4.6 non-agentic API baseline; three calls per question. | 100 |
Evidence
- page: PDF p. 13, Methods: LLM-only baselines (Default API parameters, three calls per question, no files, public prompt, and full benchmark.) — supports /benchmark_version, /scope, /protocol
- figure: PDF pp. 3–4, Figure 2A and Results (Figure labels and text report ChatGPT 5.2 at 5.3% and Claude Opus 4.6 at 3.7%.) — supports /metrics, /results
compbiobench-v1-opus-max-three-runsvv1
Evaluated models / systems: Claude Code (Opus 4.6)
Scopefull · n=100
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / effortmax
Browserweb access through the agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes initially; timed-out questions rerun clean once with 240 minutes
TemperatureNot reported
SeedNot reported
Repeats3
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticsmean over three independent runs; consistency reported as correct in all three and at least once
Contaminationsynthetic/augmented inputs and scrubbed identifiers reduce lookup shortcuts
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Accuracy | absolute | percent | mean problem accuracy over three runs | whitespace-stripped exact match |
| Wall-clock time per question | absolute | seconds | mean after 7200-second display clipping | Not reported |
| Cost per question | absolute | USD | mean after USD-10 display clipping | Not reported |
Results
Evidence
- page: PDF pp. 13–14, Methods: agent execution and model configurations (Reports Claude Code v2.1.87, claude-opus-4-6 1M, max effort, three runs, tools, and timeout policy.) — supports /benchmark_version, /scope, /protocol
- figure: PDF pp. 3–4, Figure 2A–C and Results (Reports 81.0% accuracy, 1101.0 seconds, USD 1.7, and consistency.) — supports /metrics, /results
compbiobench-v1-hardest-opus-max-three-runsvv1
Evaluated models / systems: Claude Code (Opus 4.6)
Scopesubset · n=17
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / effortmax
Browserweb access through agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes initially; timeout rerun once with 240 minutes
TemperatureNot reported
SeedNot reported
Repeats3
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticsaccuracy averaged over three runs within the 17-task subset
Contaminationbenchmark construction reduces lookup shortcuts
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Accuracy — difficulty Levels 4–5 | absolute | percent | mean problem accuracy over three runs | whitespace-stripped exact match |
Results
| Model | Metric | Value | n |
|---|
| Claude Code (Opus 4.6) | Accuracy — difficulty Levels 4–5 | 69 percent Supplementary Figure 2 label; average across three runs. | 17 |
Evidence
- page: PDF pp. 13–14, Methods (Exact Claude Code Opus configuration and common protocol.) — supports /protocol
- figure: PDF p. 11, Supplementary Figure 2B and caption (Levels 4 and 5 grouped; 17 tasks; Opus label 69; three-run average.) — supports /scope, /metrics, /results
compbiobench-v1-sonnet-high-one-runvv1
Evaluated models / systems: Claude Code (Sonnet 4.6)
Scopefull · n=100
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / efforthigh
Browserweb access through the agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes initially; timed-out questions rerun clean once with 240 minutes
TemperatureNot reported
SeedNot reported
Repeats1
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticssingle full-benchmark run
Contaminationsynthetic/augmented inputs and scrubbed identifiers reduce lookup shortcuts
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Accuracy | absolute | percent | problem-weighted mean | whitespace-stripped exact match |
| Wall-clock time per question | absolute | seconds | mean after 7200-second display clipping | Not reported |
| Cost per question | absolute | USD | mean after USD-10 display clipping | Not reported |
Results
Evidence
- page: PDF pp. 13–14, Methods (Reports Claude Code v2.1.87, claude-sonnet-4-6 1M, high effort, one run, tools, and timeout policy.) — supports /benchmark_version, /scope, /protocol
- figure: PDF pp. 3–4, Figure 2A–C (Printed labels report 70.0%, 1049.2 seconds, and USD 1.2.) — supports /metrics, /results
compbiobench-v1-hardest-sonnet-high-one-runvv1
Evaluated models / systems: Claude Code (Sonnet 4.6)
Scopesubset · n=17
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / efforthigh
Browserweb access through agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes initially; timeout rerun once with 240 minutes
TemperatureNot reported
SeedNot reported
Repeats1
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticsproblem-weighted accuracy in one run over the 17-task subset
Contaminationbenchmark construction reduces lookup shortcuts
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Accuracy — difficulty Levels 4–5 | absolute | percent | problem-weighted mean | whitespace-stripped exact match |
Results
| Model | Metric | Value | n |
|---|
| Claude Code (Sonnet 4.6) | Accuracy — difficulty Levels 4–5 | 53 percent Supplementary Figure 2 label; one run. | 17 |
Evidence
- page: PDF pp. 13–14, Methods (Exact Claude Code Sonnet configuration and common protocol.) — supports /protocol
- figure: PDF p. 11, Supplementary Figure 2B and caption (Levels 4 and 5 grouped; 17 tasks; Sonnet label 53.) — supports /scope, /metrics, /results