agentic-eval · audited · verified 2026-07-21

CompBioBench

A 100-task agent benchmark of objectively gradable computational-biology problems requiring multi-step reasoning, bespoke code, tools, and real-world external resources.

+4 more
Task, access, and protocol audit: v1 contains 100 tasks. The public TSV partitions them into 8 domains and 5 question styles; only one Structure task uses a protein PDB, while standalone binding counts are Not reported. Tasks and inputs are open, but the answer key and grader backend are private. Creator results are split by agent, effort, repeats, timeout policy, and the 17-task hardest subset.

Benchmark definition

What is counted

Version
v1
Total
100 (v1 independent computational-biology tasks)
Task formats
agentic computational analysis; exact single-line answer
Capabilities
Data analysisCodingTool useRetrievalScientific reasoning
Modalities
TextTableDNA or RNA sequenceProtein sequence3D structureRaw omicsDatabaseWebCode

Version history

VersionStatusRelease / as-ofTotalFormal tracks
v1
compbiobench-v1
current2026-04-06100 (v1 independent computational-biology tasks)None registered

Tracks and subsets

IDCountBasisPartition?Notes
Domain — Epigenomics
compbiobench-domain-epigenomics
20tasks in the domain partitionNoOfficial v1 TSV domain label; partition members are non-exhaustive here because other independent partitions are also registered.
Domain — Genomics
compbiobench-domain-genomics
20tasks in the domain partitionNoOfficial v1 TSV domain label.
Domain — Machine Learning
compbiobench-domain-machine-learning
7tasks in the domain partitionNoOfficial v1 TSV domain label.
Domain — Population Genetics
compbiobench-domain-population-genetics
12tasks in the domain partitionNoOfficial v1 TSV domain label.
Domain — Single-cell
compbiobench-domain-single-cell
21tasks in the domain partitionNoOfficial v1 TSV domain label.
Domain — Spatial
compbiobench-domain-spatial
2tasks in the domain partitionNoOfficial v1 TSV domain label.
Domain — Structure
compbiobench-domain-structure
1tasks in the domain partitionNoThe sole Structure task asks which uppercase letter a supplied PDB protein resembles across projections.
Domain — Transcriptomics
compbiobench-domain-transcriptomics
17tasks in the domain partitionNoOfficial v1 TSV domain label.
Question style — Metadata Recovery
compbiobench-style-metadata-recovery
27tasks in the question-style partitionNoOfficial v1 TSV question_style label.
Question style — Retrieval
compbiobench-style-retrieval
17tasks in the question-style partitionNoOfficial v1 TSV question_style label.
Question style — Routine Analysis
compbiobench-style-routine-analysis
22tasks in the question-style partitionNoOfficial v1 TSV question_style label.
Question style — Synthetic/Augmented Data
compbiobench-style-synthetic-augmented
26tasks in the question-style partitionNoOfficial v1 TSV question_style label.
Question style — Tooling
compbiobench-style-tooling
8tasks in the question-style partitionNoOfficial v1 TSV question_style label.
Contributor difficulty — Level 1
compbiobench-difficulty-level-1
17tasks in the paper difficulty partitionNoContributor ratings are approximate and were not extensively calibrated.
Contributor difficulty — Level 2
compbiobench-difficulty-level-2
26tasks in the paper difficulty partitionNoContributor ratings are approximate and were not extensively calibrated.
Contributor difficulty — Level 3
compbiobench-difficulty-level-3
40tasks in the paper difficulty partitionNoContributor ratings are approximate and were not extensively calibrated.
Contributor difficulty — Levels 4–5
compbiobench-difficulty-levels-4-5
17tasks in the grouped hardest subsetNoThe paper groups Levels 4 and 5 for its hardest-subset analysis.
Internet required
compbiobench-internet-required
78tasks whose v1 TSV internet_required field is TrueNoPublic v1 TSV count.
Internet not required
compbiobench-internet-not-required
22tasks whose v1 TSV internet_required field is FalseNoPublic v1 TSV count.

Scientific Task Atlas

Scientific task classification

partial for v1. Official domains are not an exhaustive scientific-task taxonomy; the runner establishes the end-to-end workflow.

Scientific taskCoverageCountMappingEvidence
End-to-end computational analysisexplicitly-in-scope100 tasks
v1 independent computational-biology tasks
official-taxonomy
high confidence
compbiobench-evidence-counts
compbiobench-evidence-runner-license
Tasks require code, tools, and external resources.
Omics and cellular analysisexplicitly-in-scopeNot reported
v1 tasks in official genomics, transcriptomics, epigenetics, single-cell, and spatial domains.
official-taxonomy
high confidence
compbiobench-evidence-taxonomy
The official domains establish broad omics coverage but do not provide an exhaustive Scientific Task breakdown.

Scientific coverage notes

DomainCoverageCountInterpretation
Protein structureobserved1The official v1 TSV contains one Structure-domain PDB projection task; this is not a protein-folding or binding benchmark.
Protein-protein bindingunknownNot reportedNo standalone protein-protein binding category or count is reported.
Protein-ligand bindingunknownNot reportedNo standalone protein-ligand binding category or count is reported.

Evaluation registry

Works and run settings

A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.

compbiobench-v1-hardest-codex-xhigh-three-runsvv1

Evaluated models / systems: Codex CLI (GPT-5.4)

Scopesubset · n=17
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / effortxhigh
Browserweb access through agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes initially; timeout rerun once with 240 minutes
TemperatureNot reported
SeedNot reported
Repeats3
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticsaccuracy averaged over three runs within the 17-task subset
Contaminationbenchmark construction reduces lookup shortcuts
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracy — difficulty Levels 4–5absolutepercentmean problem accuracy over three runswhitespace-stripped exact match

Results

ModelMetricValuen
Codex CLI (GPT-5.4)Accuracy — difficulty Levels 4–559 percent
Supplementary Figure 2 label; average across three runs.
17

Evidence

  • page: PDF pp. 13–14, Methods (Exact Codex configuration and common agent protocol.) — supports /protocol
  • figure: PDF p. 11, Supplementary Figure 2B and caption (Levels 4 and 5 grouped; 17 tasks; Codex label 59; three-run average.) — supports /scope, /metrics, /results
compbiobench-v1-codex-xhigh-three-runsvv1

Evaluated models / systems: Codex CLI (GPT-5.4)

Scopefull · n=100
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / effortxhigh
Browserweb access through the agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes initially; timed-out questions rerun clean once with 240 minutes
TemperatureNot reported
SeedNot reported
Repeats3
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticsmean over three independent runs; consistency reported as correct in all three and correct at least once
Contaminationsynthetic/augmented inputs and scrubbed identifiers reduce lookup shortcuts
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmean problem accuracy over three runswhitespace-stripped exact match
Wall-clock time per questionabsolutesecondsmean over questions and runs after 7200-second display clippingNot reported
Cost per questionabsoluteUSDmean over questions and runs after USD-10 display clippingNot reported

Results

ModelMetricValuen
Codex CLI (GPT-5.4)Accuracy83.3 percent
Mean of three full runs; 73% solved on all three and 92% at least once.
100
Codex CLI (GPT-5.4)Wall-clock time per question679 seconds
Printed Figure 2 label; mean across three runs.
100
Codex CLI (GPT-5.4)Cost per question1 USD
Printed Figure 2 label; mean across three runs.
100

Evidence

  • page: PDF pp. 13–14, Methods: agent execution and model configurations (Reports public wrapper prompt, internet/code/tool access, Conda isolation, 120/240-minute policy, Codex CLI v0.115.0, gpt-5.4, xhigh, and three runs.) — supports /benchmark_version, /scope, /protocol
  • figure: PDF pp. 3–4, Figure 2A–C and Results (Printed labels report 83.3% accuracy, 679.0 seconds, and USD 1.0; text reports three-run consistency.) — supports /metrics, /results
compbiobench-v1-haiku-one-run-120mvv1

Evaluated models / systems: Claude Code (Haiku 4.5)

Scopefull · n=100
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / effortunsupported
Browserweb access through the agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes; no 240-minute timeout rerun
TemperatureNot reported
SeedNot reported
Repeats1
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticssingle full-benchmark run
Contaminationsynthetic/augmented inputs and scrubbed identifiers reduce lookup shortcuts
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentproblem-weighted meanwhitespace-stripped exact match
Wall-clock time per questionabsolutesecondsmean after 7200-second display clippingNot reported
Cost per questionabsoluteUSDmean after USD-10 display clippingNot reported

Results

ModelMetricValuen
Claude Code (Haiku 4.5)Accuracy34 percent
One full run.
100
Claude Code (Haiku 4.5)Wall-clock time per question809.3 seconds
Printed Figure 2 label.
100
Claude Code (Haiku 4.5)Cost per question0.3 USD
Printed Figure 2 label.
100

Evidence

  • page: PDF pp. 13–14, Methods (Reports Claude Code v2.1.87, dated Haiku identifier, unsupported effort, one run, and no timeout rerun.) — supports /benchmark_version, /scope, /protocol
  • figure: PDF pp. 3–4, Figure 2A–C (Printed labels report 34.0%, 809.3 seconds, and USD 0.3.) — supports /metrics, /results
compbiobench-v1-hardest-haiku-one-run-120mvv1

Evaluated models / systems: Claude Code (Haiku 4.5)

Scopesubset · n=17
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / effortunsupported
Browserweb access through agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes; no 240-minute timeout rerun
TemperatureNot reported
SeedNot reported
Repeats1
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticsproblem-weighted accuracy in one run over the 17-task subset
Contaminationbenchmark construction reduces lookup shortcuts
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracy — difficulty Levels 4–5absolutepercentproblem-weighted meanwhitespace-stripped exact match

Results

ModelMetricValuen
Claude Code (Haiku 4.5)Accuracy — difficulty Levels 4–512 percent
Supplementary Figure 2 label; one run.
17

Evidence

  • page: PDF pp. 13–14, Methods (Exact Haiku configuration and timeout exception.) — supports /protocol
  • figure: PDF p. 11, Supplementary Figure 2B and caption (Levels 4 and 5 grouped; 17 tasks; Haiku label 12.) — supports /scope, /metrics, /results
compbiobench-v1-nonagentic-api-three-calls-no-filesvv1

Evaluated models / systems: ChatGPT 5.2, Claude Opus 4.6

Scopefull · n=100
Shotszero-shot
Turnssingle-turn API call
System prompt publicNot reported
Reasoning / effortdefault API parameters
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
Temperaturedefault
SeedNot reported
Repeats3
Graderwhitespace-stripped exact string match · human review: not reported
Statisticsmean accuracy over three calls per question
Contaminationbenchmark construction reduces lookup shortcuts
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmean over three calls per problemwhitespace-stripped exact match

Results

ModelMetricValuen
ChatGPT 5.2Accuracy5.3 percent
ChatGPT 5.2 non-agentic API baseline; three calls per question.
100
Claude Opus 4.6Accuracy3.7 percent
Claude Opus 4.6 non-agentic API baseline; three calls per question.
100

Evidence

  • page: PDF p. 13, Methods: LLM-only baselines (Default API parameters, three calls per question, no files, public prompt, and full benchmark.) — supports /benchmark_version, /scope, /protocol
  • figure: PDF pp. 3–4, Figure 2A and Results (Figure labels and text report ChatGPT 5.2 at 5.3% and Claude Opus 4.6 at 3.7%.) — supports /metrics, /results
compbiobench-v1-opus-max-three-runsvv1

Evaluated models / systems: Claude Code (Opus 4.6)

Scopefull · n=100
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / effortmax
Browserweb access through the agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes initially; timed-out questions rerun clean once with 240 minutes
TemperatureNot reported
SeedNot reported
Repeats3
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticsmean over three independent runs; consistency reported as correct in all three and at least once
Contaminationsynthetic/augmented inputs and scrubbed identifiers reduce lookup shortcuts
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmean problem accuracy over three runswhitespace-stripped exact match
Wall-clock time per questionabsolutesecondsmean after 7200-second display clippingNot reported
Cost per questionabsoluteUSDmean after USD-10 display clippingNot reported

Results

ModelMetricValuen
Claude Code (Opus 4.6)Accuracy81 percent
Mean of three full runs; 73% solved all three and 86% at least once.
100
Claude Code (Opus 4.6)Wall-clock time per question1101 seconds
Printed Figure 2 label.
100
Claude Code (Opus 4.6)Cost per question1.7 USD
Printed Figure 2 label.
100

Evidence

  • page: PDF pp. 13–14, Methods: agent execution and model configurations (Reports Claude Code v2.1.87, claude-opus-4-6 1M, max effort, three runs, tools, and timeout policy.) — supports /benchmark_version, /scope, /protocol
  • figure: PDF pp. 3–4, Figure 2A–C and Results (Reports 81.0% accuracy, 1101.0 seconds, USD 1.7, and consistency.) — supports /metrics, /results
compbiobench-v1-hardest-opus-max-three-runsvv1

Evaluated models / systems: Claude Code (Opus 4.6)

Scopesubset · n=17
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / effortmax
Browserweb access through agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes initially; timeout rerun once with 240 minutes
TemperatureNot reported
SeedNot reported
Repeats3
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticsaccuracy averaged over three runs within the 17-task subset
Contaminationbenchmark construction reduces lookup shortcuts
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracy — difficulty Levels 4–5absolutepercentmean problem accuracy over three runswhitespace-stripped exact match

Results

ModelMetricValuen
Claude Code (Opus 4.6)Accuracy — difficulty Levels 4–569 percent
Supplementary Figure 2 label; average across three runs.
17

Evidence

  • page: PDF pp. 13–14, Methods (Exact Claude Code Opus configuration and common protocol.) — supports /protocol
  • figure: PDF p. 11, Supplementary Figure 2B and caption (Levels 4 and 5 grouped; 17 tasks; Opus label 69; three-run average.) — supports /scope, /metrics, /results
compbiobench-v1-sonnet-high-one-runvv1

Evaluated models / systems: Claude Code (Sonnet 4.6)

Scopefull · n=100
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / efforthigh
Browserweb access through the agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes initially; timed-out questions rerun clean once with 240 minutes
TemperatureNot reported
SeedNot reported
Repeats1
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticssingle full-benchmark run
Contaminationsynthetic/augmented inputs and scrubbed identifiers reduce lookup shortcuts
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentproblem-weighted meanwhitespace-stripped exact match
Wall-clock time per questionabsolutesecondsmean after 7200-second display clippingNot reported
Cost per questionabsoluteUSDmean after USD-10 display clippingNot reported

Results

ModelMetricValuen
Claude Code (Sonnet 4.6)Accuracy70 percent
One full run.
100
Claude Code (Sonnet 4.6)Wall-clock time per question1049.2 seconds
Printed Figure 2 label.
100
Claude Code (Sonnet 4.6)Cost per question1.2 USD
Printed Figure 2 label.
100

Evidence

  • page: PDF pp. 13–14, Methods (Reports Claude Code v2.1.87, claude-sonnet-4-6 1M, high effort, one run, tools, and timeout policy.) — supports /benchmark_version, /scope, /protocol
  • figure: PDF pp. 3–4, Figure 2A–C (Printed labels report 70.0%, 1049.2 seconds, and USD 1.2.) — supports /metrics, /results
compbiobench-v1-hardest-sonnet-high-one-runvv1

Evaluated models / systems: Claude Code (Sonnet 4.6)

Scopesubset · n=17
Shotszero-shot
Turnsmulti-turn agent trajectory
System prompt publicNo
Reasoning / efforthigh
Browserweb access through agent environment
InternetYes
DatabasesYes
Code executionYes
ContainerNo
External toolsYes
Token budgetNot reported
Time / cost budget120 minutes initially; timeout rerun once with 240 minutes
TemperatureNot reported
SeedNot reported
Repeats1
Graderwhitespace-stripped exact string match with occasional manual format-only correction · human review: yes
Statisticsproblem-weighted accuracy in one run over the 17-task subset
Contaminationbenchmark construction reduces lookup shortcuts
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracy — difficulty Levels 4–5absolutepercentproblem-weighted meanwhitespace-stripped exact match

Results

ModelMetricValuen
Claude Code (Sonnet 4.6)Accuracy — difficulty Levels 4–553 percent
Supplementary Figure 2 label; one run.
17

Evidence

  • page: PDF pp. 13–14, Methods (Exact Claude Code Sonnet configuration and common protocol.) — supports /protocol
  • figure: PDF p. 11, Supplementary Figure 2B and caption (Levels 4 and 5 grouped; 17 tasks; Sonnet label 53.) — supports /scope, /metrics, /results

Comparable result views

Accuracy — difficulty Levels 4–5

compbiobench-codex-hardest · compbiobench-v1-hardest-codex-xhigh-three-runs

CSV ↓
Accessible data table
ModelValueComparability group
Codex CLI (GPT-5.4)59compbiobench-v1-hardest-codex-xhigh-three-runs

Accuracy

compbiobench-creator-full · compbiobench-v1-codex-xhigh-three-runs

CSV ↓
Accessible data table
ModelValueComparability group
Codex CLI (GPT-5.4)83.3compbiobench-v1-codex-xhigh-three-runs

Wall-clock time per question

compbiobench-creator-full · compbiobench-v1-codex-xhigh-three-runs

CSV ↓
Accessible data table
ModelValueComparability group
Codex CLI (GPT-5.4)679compbiobench-v1-codex-xhigh-three-runs

Cost per question

compbiobench-creator-full · compbiobench-v1-codex-xhigh-three-runs

CSV ↓
Accessible data table
ModelValueComparability group
Codex CLI (GPT-5.4)1compbiobench-v1-codex-xhigh-three-runs

Accuracy

compbiobench-haiku-full · compbiobench-v1-haiku-one-run-120m

CSV ↓
Accessible data table
ModelValueComparability group
Claude Code (Haiku 4.5)34compbiobench-v1-haiku-one-run-120m

Wall-clock time per question

compbiobench-haiku-full · compbiobench-v1-haiku-one-run-120m

CSV ↓
Accessible data table
ModelValueComparability group
Claude Code (Haiku 4.5)809.3compbiobench-v1-haiku-one-run-120m

Cost per question

compbiobench-haiku-full · compbiobench-v1-haiku-one-run-120m

CSV ↓
Accessible data table
ModelValueComparability group
Claude Code (Haiku 4.5)0.3compbiobench-v1-haiku-one-run-120m

Accuracy — difficulty Levels 4–5

compbiobench-haiku-hardest · compbiobench-v1-hardest-haiku-one-run-120m

CSV ↓
Accessible data table
ModelValueComparability group
Claude Code (Haiku 4.5)12compbiobench-v1-hardest-haiku-one-run-120m

Accuracy

compbiobench-nonagentic-baselines · compbiobench-v1-nonagentic-api-three-calls-no-files

CSV ↓
Accessible data table
ModelValueComparability group
ChatGPT 5.25.3compbiobench-v1-nonagentic-api-three-calls-no-files
Claude Opus 4.63.7compbiobench-v1-nonagentic-api-three-calls-no-files

Accuracy

compbiobench-opus-full · compbiobench-v1-opus-max-three-runs

CSV ↓
Accessible data table
ModelValueComparability group
Claude Code (Opus 4.6)81compbiobench-v1-opus-max-three-runs

Wall-clock time per question

compbiobench-opus-full · compbiobench-v1-opus-max-three-runs

CSV ↓
Accessible data table
ModelValueComparability group
Claude Code (Opus 4.6)1101compbiobench-v1-opus-max-three-runs

Cost per question

compbiobench-opus-full · compbiobench-v1-opus-max-three-runs

CSV ↓
Accessible data table
ModelValueComparability group
Claude Code (Opus 4.6)1.7compbiobench-v1-opus-max-three-runs

Accuracy — difficulty Levels 4–5

compbiobench-opus-hardest · compbiobench-v1-hardest-opus-max-three-runs

CSV ↓
Accessible data table
ModelValueComparability group
Claude Code (Opus 4.6)69compbiobench-v1-hardest-opus-max-three-runs

Accuracy

compbiobench-sonnet-full · compbiobench-v1-sonnet-high-one-run

CSV ↓
Accessible data table
ModelValueComparability group
Claude Code (Sonnet 4.6)70compbiobench-v1-sonnet-high-one-run

Wall-clock time per question

compbiobench-sonnet-full · compbiobench-v1-sonnet-high-one-run

CSV ↓
Accessible data table
ModelValueComparability group
Claude Code (Sonnet 4.6)1049.2compbiobench-v1-sonnet-high-one-run

Cost per question

compbiobench-sonnet-full · compbiobench-v1-sonnet-high-one-run

CSV ↓
Accessible data table
ModelValueComparability group
Claude Code (Sonnet 4.6)1.2compbiobench-v1-sonnet-high-one-run

Accuracy — difficulty Levels 4–5

compbiobench-sonnet-hardest · compbiobench-v1-hardest-sonnet-high-one-run

CSV ↓
Accessible data table
ModelValueComparability group
Claude Code (Sonnet 4.6)53compbiobench-v1-hardest-sonnet-high-one-run

Evidence and change history

Source locators remain visible; expand an item to inspect the exact Registry fields it supports.

Agentic systems are adept at solving well-scoped, verifiable problems in computational biology · page: pp. 1–3, abstract, benchmark design/philosophy, and Figure 1 (Defines the task format, objective-answer construction, organizations, domains, capabilities, and modalities.) · Supports 6 fields

Open source →

  • /name
  • /organizations
  • /kind
  • /summary
  • /capabilities
  • /modalities
compbiobench-zenodo-resource · release: Zenodo immutable record 19443186, published 2026-04-06 (Record title identifies CompBioBench v1 and links the concept DOI.) · Supports 4 fields

Open source →

  • /release_date
  • /latest_version
  • /resources/1/url
  • /resources/1/pin
compbiobench-zenodo-resource · repository-path: compbiobench.v1.tsv (MD5 b9d72c04c018ee25798cc93ba77c1964) (Contains 100 unique question_id rows; domain counts sum to 100, style counts sum to 100, and internet_required is 78 True / 22 False.) · Supports 8 fields

Open source →

  • /task_counts/total
  • /task_counts/basis
  • /task_counts/subsets
  • /versions/0/task_counts/total
  • /versions/0/task_counts/basis
  • /versions/0/task_counts/subsets
  • /coverage_notes/0
  • /scientific_task_classification/entries/0
compbiobench-hf-data-resource · repository-path: compbiobench.v1.tsv at commit 86ee0a22e036fef2a98f382f8fd528d5d390dde3 (Public questions, file paths, domain, style, skills, internet, and GPU flags.) · Supports 4 fields

Open source →

  • /domains
  • /coverage_notes/1
  • /coverage_notes/2
  • /scientific_task_classification/entries/1
compbiobench-leaderboard-resource · repository-path: app.py and README at commit 6a63e6d2cae531d9e9bf46b341606ed6b304ba3f (Public exact-match submission UI backed by private submissions, ground-truth, and results repositories.) · Supports 6 fields

Open source →

  • /access/level
  • /access/grader
  • /access/license
  • /resources/4/license
  • /resources/4/pin
  • /implementations/1
Agentic systems are adept at solving well-scoped, verifiable problems in computational biology · page: PDF page headers, first-page copyright statement, and Data availability (Preprint is CC BY-NC 4.0 and links the official data, runner, and leaderboard resources.) · Supports 4 fields

Open source →

  • /access/license
  • /resources
  • /resources/0/license
  • /resources/0/pin
compbiobench-zenodo-resource · release: Zenodo record 19443186 metadata and file manifest (Dataset is CC BY 4.0 and provides the immutable TSV/data-archive files.) · Supports 5 fields

Open source →

  • /access/tasks
  • /access/artifacts
  • /access/license
  • /resources/1/license
  • /resources/2/license
compbiobench-runner-resource · repository-path: LICENSE, README.md, run_benchmark.py, and environment.yml at dc350ed37ccd7d7ce96347d139f06dc4bf283f26 (MIT license and per-question Conda execution without grading.) · Supports 5 fields

Open source →

  • /access/license
  • /resources/3/license
  • /resources/3/pin
  • /implementations/0
  • /scientific_task_classification/entries/0

View source-level modification history on GitHub →