agentic-eval · audited-with-caveats · verified 2026-07-21

GeneBench-Pro

A research-level agent benchmark of 129 synthetic, multistage computational-biology analyses that require iterative QC, statistical modeling, diagnostics, and decision-relevant judgment.

+5 more
Audited with caveats: 1 field(s) are marked provisional or conflicted. Warnings are shown next to affected values and these claims are excluded from unqualified comparisons.
Release and protocol audit: the 129-problem suite is partitioned into 10 public, 50 Artificial Analysis, and 69 internal-holdout problems. All 60 reported model configurations are normalized across 13 effort/repeat groups: standard runs use 10 attempts per problem, while Claude Opus and GPT Pro (Extended) use 5. The public package's dataset card says CC BY 4.0, but its root LICENSE says MIT, so the license remains Conflicted.

Benchmark definition

What is counted

Version
paper-v1
Total
129 (self-contained synthetic scientific-analysis problems (called evaluations in the paper abstract))
Task formats
isolated-workspace scientific analysis; exact JSON response
Capabilities
Data analysisCodingTool useScientific reasoning
Modalities
TextTableRaw omicsCode

Version history

VersionStatusRelease / as-ofTotalFormal tracks
paper-v1
genebench-pro-paper-v1
current2026-06-30129 (self-contained synthetic scientific-analysis problems (called evaluations in the paper abstract))None registered

Tracks and subsets

IDCountBasisPartition?Notes
Public release subset
genebench-pro-public-release
10problemsExclusive & exhaustiveTen externally reviewed case-study packages are public with prompts, staged data, ground truth, grader configuration, and reports.
Artificial Analysis reporting subset
genebench-pro-artificial-analysis
50held-out problemsExclusive & exhaustiveDisjoint from the public release. The formal paper says this locked reporting subset was provided to Artificial Analysis; the launch page still describes delivery in future tense. The tasks are not public.
Internal holdout
genebench-pro-internal-holdout
69held-out problemsExclusive & exhaustiveRemainder after the disjoint 10-problem public and 50-problem Artificial Analysis subsets; not publicly released.
Primary domain — Statistical genetics
genebench-pro-primary-statistical-genetics
17problems in the primary-domain atlasNoOne member of the 10-domain partition; marked non-exhaustive here because the registry also stores independent release and review partitions.
Primary domain — Population genetics
genebench-pro-primary-population-genetics
21problems in the primary-domain atlasNoOne member of the 10-domain partition; marked non-exhaustive here because the registry also stores independent release and review partitions.
Primary domain — Quantitative genetics
genebench-pro-primary-quantitative-genetics
17problems in the primary-domain atlasNoOne member of the 10-domain partition; marked non-exhaustive here because the registry also stores independent release and review partitions.
Primary domain — Regulatory omics
genebench-pro-primary-regulatory-omics
17problems in the primary-domain atlasNoOne member of the 10-domain partition; marked non-exhaustive here because the registry also stores independent release and review partitions.
Primary domain — Functional genomics
genebench-pro-primary-functional-genomics
9problems in the primary-domain atlasNoOne member of the 10-domain partition; marked non-exhaustive here because the registry also stores independent release and review partitions.
Primary domain — Proteomics
genebench-pro-primary-proteomics
7problems in the primary-domain atlasNoOne member of the 10-domain partition; marked non-exhaustive here because the registry also stores independent release and review partitions.
Primary domain — Clinical, PGx & diagnostics
genebench-pro-primary-clinical-pgx-diagnostics
26problems in the primary-domain atlasNoOne member of the 10-domain partition; marked non-exhaustive here because the registry also stores independent release and review partitions.
Primary domain — Cancer genomics
genebench-pro-primary-cancer-genomics
10problems in the primary-domain atlasNoOne member of the 10-domain partition; marked non-exhaustive here because the registry also stores independent release and review partitions.
Primary domain — Microbial genomics
genebench-pro-primary-microbial-genomics
3problems in the primary-domain atlasNoOne member of the 10-domain partition; marked non-exhaustive here because the registry also stores independent release and review partitions.
Primary domain — Forensic genetics
genebench-pro-primary-forensic-genetics
2problems in the primary-domain atlasNoOne member of the 10-domain partition; marked non-exhaustive here because the registry also stores independent release and review partitions.
Terminal subdomain — Association & correction
genebench-pro-terminal-association-correction
6problems in the terminal-subdomain atlasNoOne member of the 21-terminal-subdomain partition; stored alongside independent release and review partitions.
Terminal subdomain — Causal mapping
genebench-pro-terminal-causal-mapping
6problems in the terminal-subdomain atlasNoOne member of the 21-terminal-subdomain partition; stored alongside independent release and review partitions.
Terminal subdomain — Heritability and architecture
genebench-pro-terminal-heritability-architecture
2problems in the terminal-subdomain atlasNoOne member of the 21-terminal-subdomain partition; stored alongside independent release and review partitions.
Terminal subdomain — Pedigree, IBD, and phasing
genebench-pro-terminal-pedigree-ibd-phasing
3problems in the terminal-subdomain atlasNoOne member of the 21-terminal-subdomain partition; stored alongside independent release and review partitions.
Terminal subdomain — Selection & mutation
genebench-pro-terminal-selection-mutation
7problems in the terminal-subdomain atlasNoOne member of the 21-terminal-subdomain partition; stored alongside independent release and review partitions.
Terminal subdomain — Admixture & ancient DNA
genebench-pro-terminal-admixture-adna
6problems in the terminal-subdomain atlasNoOne member of the 21-terminal-subdomain partition; stored alongside independent release and review partitions.
Terminal subdomain — History & genealogies
genebench-pro-terminal-history-genealogies
8problems in the terminal-subdomain atlasNoOne member of the 21-terminal-subdomain partition; stored alongside independent release and review partitions.
Terminal subdomain — Trait architecture and variance
genebench-pro-terminal-trait-architecture-variance
6problems in the terminal-subdomain atlasNoOne member of the 21-terminal-subdomain partition; stored alongside independent release and review partitions.
Terminal subdomain — Family, social, and transmission effects
genebench-pro-terminal-family-social-transmission
6problems in the terminal-subdomain atlasNoOne member of the 21-terminal-subdomain partition; stored alongside independent release and review partitions.
Terminal subdomain — Polygenic prediction and genomic selection
genebench-pro-terminal-polygenic-prediction-selection
5problems in the terminal-subdomain atlasNoOne member of the 21-terminal-subdomain partition; stored alongside independent release and review partitions.
Terminal subdomain — Regulatory QTLs & ASE
genebench-pro-terminal-regulatory-qtls-ase
8problems in the terminal-subdomain atlasNoOne member of the 21-terminal-subdomain partition; stored alongside independent release and review partitions.
Terminal subdomain — Transcriptome structure
genebench-pro-terminal-transcriptome-structure
5problems in the terminal-subdomain atlasNoOne member of the 21-terminal-subdomain partition; stored alongside independent release and review partitions.
Terminal subdomain — Spatial and chromatin context
genebench-pro-terminal-spatial-chromatin-context
4problems in the terminal-subdomain atlasNoOne member of the 21-terminal-subdomain partition; stored alongside independent release and review partitions.
Terminal subdomain — Functional genomics
genebench-pro-terminal-functional-genomics-terminal
9problems in the terminal-subdomain atlasNoOne member of the 21-terminal-subdomain partition; stored alongside independent release and review partitions.
Terminal subdomain — Proteomics and biomarkers
genebench-pro-terminal-proteomics-biomarkers
7problems in the terminal-subdomain atlasNoOne member of the 21-terminal-subdomain partition; stored alongside independent release and review partitions.
Terminal subdomain — Clinical variant interpretation & penetrance
genebench-pro-terminal-clinical-variant-penetrance
11problems in the terminal-subdomain atlasNoOne member of the 21-terminal-subdomain partition; stored alongside independent release and review partitions.
Terminal subdomain — Pharmacogenomics and treatment response
genebench-pro-terminal-pharmacogenomics-treatment-response
8problems in the terminal-subdomain atlasNoOne member of the 21-terminal-subdomain partition; stored alongside independent release and review partitions.
Terminal subdomain — Prenatal, reproductive, and clinical-risk genetics
genebench-pro-terminal-prenatal-reproductive-risk
7problems in the terminal-subdomain atlasNoOne member of the 21-terminal-subdomain partition; stored alongside independent release and review partitions.
Terminal subdomain — Cancer somatic genomics and liquid biopsy
genebench-pro-terminal-cancer-somatic-liquid-biopsy
10problems in the terminal-subdomain atlasNoOne member of the 21-terminal-subdomain partition; stored alongside independent release and review partitions.
Terminal subdomain — Microbial and metagenomic genomics
genebench-pro-terminal-microbial-metagenomic-genomics
3problems in the terminal-subdomain atlasNoOne member of the 21-terminal-subdomain partition; stored alongside independent release and review partitions.
Terminal subdomain — Forensic genetics
genebench-pro-terminal-forensic-genetics-terminal
2problems in the terminal-subdomain atlasNoOne member of the 21-terminal-subdomain partition; stored alongside independent release and review partitions.
Externally reviewed problems
genebench-pro-externally-reviewed
82problemsNoOrthogonal review-status stratum; all ten public problems are included.
Not externally reviewed
genebench-pro-not-externally-reviewed
47problemsNoOrthogonal review-status stratum.

Scientific Task Atlas

Scientific task classification

partial for paper-v1. Official primary domains and task styles are not an exhaustive scientific-task taxonomy.

Scientific taskCoverageCountMappingEvidence
End-to-end computational analysisexplicitly-in-scope129 problems
self-contained synthetic scientific-analysis problems (called evaluations in the paper abstract)
official-taxonomy
high confidence
genebench-pro-paper-evidence
Problems require multistage analysis, diagnostics, and judgment.
Omics and cellular analysisexplicitly-in-scopeNot reported
Problems assigned to transcriptomics, epigenomics, single-cell, spatial, proteomics, microbiome, and related official domains.
official-taxonomy
high confidence
genebench-pro-paper-evidence
The official domain atlas supports broad omics coverage but not an exhaustive leaf-task subtotal.

Scientific coverage notes

DomainCoverageCountInterpretation
Proteomicsexplicitly-in-scope7Seven primary-domain problems are assigned to Proteomics; proteomics also appears in cross-domain problems.
Protein-protein bindingunknownNot reportedThe domain atlas covers proteomics and biomarkers but does not define or count protein-binding problems separately.

Evaluation registry

Works and run settings

A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.

genebench-pro-paper-v1-full-claude-highvpaper-v1

Evaluated models / systems: Claude Opus 4.8

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortprovider-reported high reasoning setting
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats5
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
Claude Opus 4.8Eval-level pass rate9 percent
Supplementary Table 1 configuration ClaudeOpus4.8 (high); nominal attempts per problem 5, valid-attempt mean 5.0 and range 4–5; average tokens not reported; problem regimes 0%=78.3%, 0–10%=0.0%, 10–50%=14.7%, ≥50%=7.0%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-claude-lowvpaper-v1

Evaluated models / systems: Claude Opus 4.8

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortprovider-reported low reasoning setting
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats5
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
Claude Opus 4.8Eval-level pass rate4.3 percent
Supplementary Table 1 configuration ClaudeOpus4.8 (low); nominal attempts per problem 5, valid-attempt mean 5.0 and range 3–5; average tokens not reported; problem regimes 0%=88.4%, 0–10%=0.0%, 10–50%=10.1%, ≥50%=1.6%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-claude-maxvpaper-v1

Evaluated models / systems: Claude Opus 4.8

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortprovider-reported max reasoning setting
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats5
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
Claude Opus 4.8Eval-level pass rate16 percent
Supplementary Table 1 configuration ClaudeOpus4.8 (max); nominal attempts per problem 5, valid-attempt mean 5.0 and range 5–5; average tokens not reported; problem regimes 0%=67.4%, 0–10%=0.0%, 10–50%=18.6%, ≥50%=14.0%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-claude-mediumvpaper-v1

Evaluated models / systems: Claude Opus 4.8

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortprovider-reported medium reasoning setting
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats5
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
Claude Opus 4.8Eval-level pass rate4.3 percent
Supplementary Table 1 configuration ClaudeOpus4.8 (medium); nominal attempts per problem 5, valid-attempt mean 5.0 and range 5–5; average tokens not reported; problem regimes 0%=87.6%, 0–10%=0.0%, 10–50%=10.1%, ≥50%=2.3%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-claude-xhighvpaper-v1

Evaluated models / systems: Claude Opus 4.8

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortprovider-reported xhigh reasoning setting
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats5
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
Claude Opus 4.8Eval-level pass rate10.1 percent
Supplementary Table 1 configuration ClaudeOpus4.8 (xhigh); nominal attempts per problem 5, valid-attempt mean 5.0 and range 4–5; average tokens not reported; problem regimes 0%=75.2%, 0–10%=0.0%, 10–50%=17.1%, ≥50%=7.8%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-standard-xhighvpaper-v1

Evaluated models / systems: DeepSeek V4 Flash, DeepSeek V4 Pro, GLM 5.1, GPT-5.2, GPT-5.4, GPT-5.5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, Kimi K2.6, MiMo V2.5, MiMo V2.5 Pro, MiniMax M2.7, MiniMax M3, Qwen 3.7 Max, Qwen 3.7 Plus, Tencent HY 3 Preview

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortprovider-reported xhigh reasoning setting
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats10
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
MiniMax M2.7Eval-level pass rate0.6 percent
Supplementary Table 1 configuration MiniMaxM2.7 (xhigh); nominal attempts per problem 10, valid-attempt mean 9.9 and range 8–10; average tokens not reported; problem regimes 0%=96.1%, 0–10%=2.3%, 10–50%=1.6%, ≥50%=0.0%.
129
Tencent HY 3 PreviewEval-level pass rate0.9 percent
Supplementary Table 1 configuration TencentHY3Preview (xhigh); nominal attempts per problem 10, valid-attempt mean 9.8 and range 8–10; average tokens not reported; problem regimes 0%=92.2%, 0–10%=6.2%, 10–50%=1.6%, ≥50%=0.0%.
129
MiniMax M3Eval-level pass rate0.9 percent
Supplementary Table 1 configuration MiniMaxM3 (xhigh); nominal attempts per problem 10, valid-attempt mean 9.6 and range 7–10; average tokens not reported; problem regimes 0%=96.1%, 0–10%=1.6%, 10–50%=2.3%, ≥50%=0.0%.
129
MiMo V2.5Eval-level pass rate1.2 percent
Supplementary Table 1 configuration MiMoV2.5 (xhigh); nominal attempts per problem 10, valid-attempt mean 9.9 and range 8–10; average tokens not reported; problem regimes 0%=91.5%, 0–10%=6.2%, 10–50%=2.3%, ≥50%=0.0%.
129
GLM 5.1Eval-level pass rate1.2 percent
Supplementary Table 1 configuration GLM5.1 (xhigh); nominal attempts per problem 10, valid-attempt mean 9.7 and range 8–10; average tokens not reported; problem regimes 0%=95.3%, 0–10%=2.3%, 10–50%=1.6%, ≥50%=0.8%.
129
MiMo V2.5 ProEval-level pass rate2 percent
Supplementary Table 1 configuration MiMoV2.5Pro (xhigh); nominal attempts per problem 10, valid-attempt mean 9.9 and range 8–10; average tokens not reported; problem regimes 0%=86.8%, 0–10%=9.3%, 10–50%=3.1%, ≥50%=0.8%.
129
Qwen 3.7 PlusEval-level pass rate2.3 percent
Supplementary Table 1 configuration Qwen3.7Plus (xhigh); nominal attempts per problem 10, valid-attempt mean 9.7 and range 6–10; average tokens not reported; problem regimes 0%=86.8%, 0–10%=7.0%, 10–50%=5.4%, ≥50%=0.8%.
129
DeepSeek V4 FlashEval-level pass rate2.4 percent
Supplementary Table 1 configuration DeepSeekV4Flash (xhigh); nominal attempts per problem 10, valid-attempt mean 9.9 and range 9–10; average tokens not reported; problem regimes 0%=87.6%, 0–10%=6.2%, 10–50%=4.7%, ≥50%=1.6%.
129
DeepSeek V4 ProEval-level pass rate2.4 percent
Supplementary Table 1 configuration DeepSeekV4Pro (xhigh); nominal attempts per problem 10, valid-attempt mean 9.6 and range 8–10; average tokens not reported; problem regimes 0%=84.5%, 0–10%=4.7%, 10–50%=10.9%, ≥50%=0.0%.
129
Qwen 3.7 MaxEval-level pass rate4 percent
Supplementary Table 1 configuration Qwen3.7Max (xhigh); nominal attempts per problem 10, valid-attempt mean 9.8 and range 9–10; average tokens not reported; problem regimes 0%=79.8%, 0–10%=7.8%, 10–50%=11.6%, ≥50%=0.8%.
129
Kimi K2.6Eval-level pass rate4.4 percent
Supplementary Table 1 configuration KimiK2.6 (xhigh); nominal attempts per problem 10, valid-attempt mean 9.9 and range 7–10; average tokens not reported; problem regimes 0%=84.5%, 0–10%=6.2%, 10–50%=6.2%, ≥50%=3.1%.
129
GPT-5.2Eval-level pass rate4.9 percent
Supplementary Table 1 configuration GPT-5.2 (xhigh); nominal attempts per problem 10, valid-attempt mean 9.9 and range 9–10; average tokens 52.0k; problem regimes 0%=77.5%, 0–10%=10.1%, 10–50%=10.9%, ≥50%=1.6%.
129
GPT-5.4Eval-level pass rate8.9 percent
Supplementary Table 1 configuration GPT-5.4 (xhigh); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 44.7k; problem regimes 0%=67.4%, 0–10%=9.3%, 10–50%=18.6%, ≥50%=4.7%.
129
GPT-5.5Eval-level pass rate12 percent
Supplementary Table 1 configuration GPT-5.5 (xhigh); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 28.7k; problem regimes 0%=64.3%, 0–10%=8.5%, 10–50%=18.6%, ≥50%=8.5%.
129
GPT-5.6 LunaEval-level pass rate10.8 percent
Supplementary Table 1 configuration GPT-5.6Luna (xhigh); nominal attempts per problem 10, valid-attempt mean 9.8 and range 7–10; average tokens 53.1k; problem regimes 0%=70.5%, 0–10%=8.5%, 10–50%=10.1%, ≥50%=10.9%.
129
GPT-5.6 TerraEval-level pass rate18.8 percent
Supplementary Table 1 configuration GPT-5.6Terra (xhigh); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 31.1k; problem regimes 0%=56.6%, 0–10%=7.8%, 10–50%=19.4%, ≥50%=16.3%.
129
GPT-5.6 SolEval-level pass rate26.8 percent
Supplementary Table 1 configuration GPT-5.6Sol (xhigh); nominal attempts per problem 10, valid-attempt mean 9.9 and range 9–10; average tokens 25.7k; problem regimes 0%=50.4%, 0–10%=6.2%, 10–50%=14.0%, ≥50%=29.5%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-pro-extendedvpaper-v1

Evaluated models / systems: GPT-5.2 Pro (Extended), GPT-5.4 Pro (Extended), GPT-5.5 Pro (Extended), GPT-5.6 Luna Pro (Extended), GPT-5.6 Sol Pro (Extended), GPT-5.6 Terra Pro (Extended)

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortGPT Pro (Extended)
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats5
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
GPT-5.2 Pro (Extended)Eval-level pass rate8.5 percent
Supplementary Table 1 configuration GPT-5.2Pro (Extended); nominal attempts per problem 5, valid-attempt mean 5.0 and range 4–5; average tokens not reported; problem regimes 0%=79.1%, 0–10%=0.0%, 10–50%=14.0%, ≥50%=7.0%.
129
GPT-5.4 Pro (Extended)Eval-level pass rate16.3 percent
Supplementary Table 1 configuration GPT-5.4Pro (Extended); nominal attempts per problem 5, valid-attempt mean 5.0 and range 5–5; average tokens not reported; problem regimes 0%=72.1%, 0–10%=0.0%, 10–50%=12.4%, ≥50%=15.5%.
129
GPT-5.5 Pro (Extended)Eval-level pass rate20.5 percent
Supplementary Table 1 configuration GPT-5.5Pro (Extended); nominal attempts per problem 5, valid-attempt mean 5.0 and range 5–5; average tokens not reported; problem regimes 0%=66.7%, 0–10%=0.0%, 10–50%=14.0%, ≥50%=19.4%.
129
GPT-5.6 Luna Pro (Extended)Eval-level pass rate23.6 percent
Supplementary Table 1 configuration GPT-5.6LunaPro (Extended); nominal attempts per problem 5, valid-attempt mean 5.0 and range 5–5; average tokens not reported; problem regimes 0%=65.9%, 0–10%=0.0%, 10–50%=10.9%, ≥50%=23.3%.
129
GPT-5.6 Terra Pro (Extended)Eval-level pass rate28.5 percent
Supplementary Table 1 configuration GPT-5.6TerraPro (Extended); nominal attempts per problem 5, valid-attempt mean 5.0 and range 5–5; average tokens not reported; problem regimes 0%=59.7%, 0–10%=0.0%, 10–50%=11.6%, ≥50%=28.7%.
129
GPT-5.6 Sol Pro (Extended)Eval-level pass rate31.5 percent
Supplementary Table 1 configuration GPT-5.6SolPro (Extended); nominal attempts per problem 5, valid-attempt mean 5.0 and range 5–5; average tokens not reported; problem regimes 0%=55.8%, 0–10%=0.0%, 10–50%=14.0%, ≥50%=30.2%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-reasoning-enabledvpaper-v1

Evaluated models / systems: Kimi K2.7 Code

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortreasoning_enabled
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats10
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
Kimi K2.7 CodeEval-level pass rate2.3 percent
Supplementary Table 1 configuration KimiK2.7Code (reasoning_enabled); nominal attempts per problem 10, valid-attempt mean 9.9 and range 9–10; average tokens not reported; problem regimes 0%=84.5%, 0–10%=9.3%, 10–50%=6.2%, ≥50%=0.0%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-standard-highvpaper-v1

Evaluated models / systems: Gemini 3.1 Pro, Gemini 3.5 Flash, GLM 5.2, GPT-5.2, GPT-5.4, GPT-5.5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, Grok 4.3

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortprovider-reported high reasoning setting
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats10
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
Grok 4.3Eval-level pass rate1.5 percent
Supplementary Table 1 configuration Grok4.3 (high); nominal attempts per problem 10, valid-attempt mean 10.0 and range 10–10; average tokens not reported; problem regimes 0%=92.2%, 0–10%=4.7%, 10–50%=2.3%, ≥50%=0.8%.
129
Gemini 3.1 ProEval-level pass rate3.1 percent
Supplementary Table 1 configuration Gemini3.1Pro (high); nominal attempts per problem 10, valid-attempt mean 10.0 and range 10–10; average tokens not reported; problem regimes 0%=81.4%, 0–10%=13.2%, 10–50%=4.7%, ≥50%=0.8%.
129
GLM 5.2Eval-level pass rate4.6 percent
Supplementary Table 1 configuration GLM5.2 (high); nominal attempts per problem 10, valid-attempt mean 9.8 and range 8–10; average tokens not reported; problem regimes 0%=77.5%, 0–10%=10.1%, 10–50%=10.9%, ≥50%=1.6%.
129
Gemini 3.5 FlashEval-level pass rate8.1 percent
Supplementary Table 1 configuration Gemini3.5Flash (high); nominal attempts per problem 10, valid-attempt mean 10.0 and range 10–10; average tokens not reported; problem regimes 0%=70.5%, 0–10%=14.7%, 10–50%=9.3%, ≥50%=5.4%.
129
GPT-5.2Eval-level pass rate3.5 percent
Supplementary Table 1 configuration GPT-5.2 (high); nominal attempts per problem 10, valid-attempt mean 9.9 and range 9–10; average tokens 23.2k; problem regimes 0%=83.7%, 0–10%=7.8%, 10–50%=7.0%, ≥50%=1.6%.
129
GPT-5.4Eval-level pass rate7 percent
Supplementary Table 1 configuration GPT-5.4 (high); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 27.3k; problem regimes 0%=70.5%, 0–10%=11.6%, 10–50%=13.2%, ≥50%=4.7%.
129
GPT-5.5Eval-level pass rate9.3 percent
Supplementary Table 1 configuration GPT-5.5 (high); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 19.8k; problem regimes 0%=70.5%, 0–10%=11.6%, 10–50%=10.9%, ≥50%=7.0%.
129
GPT-5.6 LunaEval-level pass rate8 percent
Supplementary Table 1 configuration GPT-5.6Luna (high); nominal attempts per problem 10, valid-attempt mean 9.8 and range 7–10; average tokens 32.3k; problem regimes 0%=76.7%, 0–10%=8.5%, 10–50%=7.8%, ≥50%=7.0%.
129
GPT-5.6 TerraEval-level pass rate16.2 percent
Supplementary Table 1 configuration GPT-5.6Terra (high); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 22.2k; problem regimes 0%=59.7%, 0–10%=12.4%, 10–50%=11.6%, ≥50%=16.3%.
129
GPT-5.6 SolEval-level pass rate24.4 percent
Supplementary Table 1 configuration GPT-5.6Sol (high); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 19.5k; problem regimes 0%=51.9%, 0–10%=7.8%, 10–50%=17.8%, ≥50%=22.5%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-standard-lowvpaper-v1

Evaluated models / systems: GPT-5.2, GPT-5.4, GPT-5.5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortprovider-reported low reasoning setting
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats10
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
GPT-5.2Eval-level pass rate1.1 percent
Supplementary Table 1 configuration GPT-5.2 (low); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 7.0k; problem regimes 0%=91.5%, 0–10%=6.2%, 10–50%=2.3%, ≥50%=0.0%.
129
GPT-5.4Eval-level pass rate3 percent
Supplementary Table 1 configuration GPT-5.4 (low); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 8.2k; problem regimes 0%=85.3%, 0–10%=7.0%, 10–50%=7.0%, ≥50%=0.8%.
129
GPT-5.5Eval-level pass rate2.4 percent
Supplementary Table 1 configuration GPT-5.5 (low); nominal attempts per problem 10, valid-attempt mean 9.9 and range 8–10; average tokens 2.8k; problem regimes 0%=87.6%, 0–10%=6.2%, 10–50%=4.7%, ≥50%=1.6%.
129
GPT-5.6 LunaEval-level pass rate2.3 percent
Supplementary Table 1 configuration GPT-5.6Luna (low); nominal attempts per problem 10, valid-attempt mean 10.0 and range 10–10; average tokens 3.6k; problem regimes 0%=89.1%, 0–10%=5.4%, 10–50%=4.7%, ≥50%=0.8%.
129
GPT-5.6 TerraEval-level pass rate6.5 percent
Supplementary Table 1 configuration GPT-5.6Terra (low); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 5.5k; problem regimes 0%=75.2%, 0–10%=10.1%, 10–50%=10.9%, ≥50%=3.9%.
129
GPT-5.6 SolEval-level pass rate14.4 percent
Supplementary Table 1 configuration GPT-5.6Sol (low); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 5.6k; problem regimes 0%=58.9%, 0–10%=12.4%, 10–50%=17.1%, ≥50%=11.6%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-standard-maxvpaper-v1

Evaluated models / systems: GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortprovider-reported max reasoning setting
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats10
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
GPT-5.6 LunaEval-level pass rate16.5 percent
Supplementary Table 1 configuration GPT-5.6Luna (max); nominal attempts per problem 10, valid-attempt mean 9.6 and range 2–10; average tokens 118.2k; problem regimes 0%=64.3%, 0–10%=4.7%, 10–50%=16.3%, ≥50%=14.7%.
129
GPT-5.6 TerraEval-level pass rate23.3 percent
Supplementary Table 1 configuration GPT-5.6Terra (max); nominal attempts per problem 10, valid-attempt mean 10.0 and range 10–10; average tokens 54.3k; problem regimes 0%=49.6%, 0–10%=9.3%, 10–50%=19.4%, ≥50%=21.7%.
129
GPT-5.6 SolEval-level pass rate28.7 percent
Supplementary Table 1 configuration GPT-5.6Sol (max); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 33.2k; problem regimes 0%=45.7%, 0–10%=10.1%, 10–50%=14.0%, ≥50%=30.2%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-standard-mediumvpaper-v1

Evaluated models / systems: GPT-5.2, GPT-5.4, GPT-5.5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortprovider-reported medium reasoning setting
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats10
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
GPT-5.2Eval-level pass rate2.4 percent
Supplementary Table 1 configuration GPT-5.2 (medium); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 18.1k; problem regimes 0%=86.0%, 0–10%=10.1%, 10–50%=3.1%, ≥50%=0.8%.
129
GPT-5.4Eval-level pass rate5 percent
Supplementary Table 1 configuration GPT-5.4 (medium); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 18.6k; problem regimes 0%=74.4%, 0–10%=17.1%, 10–50%=7.0%, ≥50%=1.6%.
129
GPT-5.5Eval-level pass rate5.9 percent
Supplementary Table 1 configuration GPT-5.5 (medium); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 10.6k; problem regimes 0%=79.1%, 0–10%=7.8%, 10–50%=9.3%, ≥50%=3.9%.
129
GPT-5.6 LunaEval-level pass rate4.7 percent
Supplementary Table 1 configuration GPT-5.6Luna (medium); nominal attempts per problem 10, valid-attempt mean 9.9 and range 8–10; average tokens 15.6k; problem regimes 0%=83.7%, 0–10%=7.0%, 10–50%=4.7%, ≥50%=4.7%.
129
GPT-5.6 TerraEval-level pass rate13.6 percent
Supplementary Table 1 configuration GPT-5.6Terra (medium); nominal attempts per problem 10, valid-attempt mean 10.0 and range 10–10; average tokens 15.9k; problem regimes 0%=65.9%, 0–10%=8.5%, 10–50%=12.4%, ≥50%=13.2%.
129
GPT-5.6 SolEval-level pass rate22.5 percent
Supplementary Table 1 configuration GPT-5.6Sol (medium); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 14.4k; problem regimes 0%=53.5%, 0–10%=7.8%, 10–50%=16.3%, ≥50%=22.5%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-standard-nonevpaper-v1

Evaluated models / systems: GPT-5.2, GPT-5.4, GPT-5.5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortprovider-reported none reasoning setting
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats10
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
GPT-5.2Eval-level pass rate0.5 percent
Supplementary Table 1 configuration GPT-5.2 (none); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 1.1k; problem regimes 0%=96.9%, 0–10%=2.3%, 10–50%=0.8%, ≥50%=0.0%.
129
GPT-5.4Eval-level pass rate0.9 percent
Supplementary Table 1 configuration GPT-5.4 (none); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 2.1k; problem regimes 0%=94.6%, 0–10%=3.9%, 10–50%=1.6%, ≥50%=0.0%.
129
GPT-5.5Eval-level pass rate0.8 percent
Supplementary Table 1 configuration GPT-5.5 (none); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 1.1k; problem regimes 0%=96.1%, 0–10%=2.3%, 10–50%=1.6%, ≥50%=0.0%.
129
GPT-5.6 LunaEval-level pass rate0.8 percent
Supplementary Table 1 configuration GPT-5.6Luna (none); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 975; problem regimes 0%=93.8%, 0–10%=4.7%, 10–50%=1.6%, ≥50%=0.0%.
129
GPT-5.6 TerraEval-level pass rate1 percent
Supplementary Table 1 configuration GPT-5.6Terra (none); nominal attempts per problem 10, valid-attempt mean 10.0 and range 10–10; average tokens 930; problem regimes 0%=95.3%, 0–10%=3.1%, 10–50%=0.8%, ≥50%=0.8%.
129
GPT-5.6 SolEval-level pass rate3.7 percent
Supplementary Table 1 configuration GPT-5.6Sol (none); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 1.4k; problem regimes 0%=82.2%, 0–10%=9.3%, 10–50%=7.0%, ≥50%=1.6%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results

Comparable result views

Eval-level pass rate

genebench-pro-claude-high · genebench-pro-paper-v1-full-claude-high

CSV ↓
Accessible data table
ModelValueComparability group
Claude Opus 4.89genebench-pro-paper-v1-full-claude-high

Eval-level pass rate

genebench-pro-claude-low · genebench-pro-paper-v1-full-claude-low

CSV ↓
Accessible data table
ModelValueComparability group
Claude Opus 4.84.3genebench-pro-paper-v1-full-claude-low

Eval-level pass rate

genebench-pro-claude-max · genebench-pro-paper-v1-full-claude-max

CSV ↓
Accessible data table
ModelValueComparability group
Claude Opus 4.816genebench-pro-paper-v1-full-claude-max

Eval-level pass rate

genebench-pro-claude-medium · genebench-pro-paper-v1-full-claude-medium

CSV ↓
Accessible data table
ModelValueComparability group
Claude Opus 4.84.3genebench-pro-paper-v1-full-claude-medium

Eval-level pass rate

genebench-pro-claude-xhigh · genebench-pro-paper-v1-full-claude-xhigh

CSV ↓
Accessible data table
ModelValueComparability group
Claude Opus 4.810.1genebench-pro-paper-v1-full-claude-xhigh

Eval-level pass rate

genebench-pro-official · genebench-pro-paper-v1-full-standard-xhigh

CSV ↓
Accessible data table
ModelValueComparability group
MiniMax M2.70.6genebench-pro-paper-v1-full-standard-xhigh
Tencent HY 3 Preview0.9genebench-pro-paper-v1-full-standard-xhigh
MiniMax M30.9genebench-pro-paper-v1-full-standard-xhigh
MiMo V2.51.2genebench-pro-paper-v1-full-standard-xhigh
GLM 5.11.2genebench-pro-paper-v1-full-standard-xhigh
MiMo V2.5 Pro2genebench-pro-paper-v1-full-standard-xhigh
Qwen 3.7 Plus2.3genebench-pro-paper-v1-full-standard-xhigh
DeepSeek V4 Flash2.4genebench-pro-paper-v1-full-standard-xhigh
DeepSeek V4 Pro2.4genebench-pro-paper-v1-full-standard-xhigh
Qwen 3.7 Max4genebench-pro-paper-v1-full-standard-xhigh
Kimi K2.64.4genebench-pro-paper-v1-full-standard-xhigh
GPT-5.24.9genebench-pro-paper-v1-full-standard-xhigh
GPT-5.48.9genebench-pro-paper-v1-full-standard-xhigh
GPT-5.512genebench-pro-paper-v1-full-standard-xhigh
GPT-5.6 Luna10.8genebench-pro-paper-v1-full-standard-xhigh
GPT-5.6 Terra18.8genebench-pro-paper-v1-full-standard-xhigh
GPT-5.6 Sol26.8genebench-pro-paper-v1-full-standard-xhigh

Eval-level pass rate

genebench-pro-pro-mode · genebench-pro-paper-v1-full-pro-extended

CSV ↓
Accessible data table
ModelValueComparability group
GPT-5.2 Pro (Extended)8.5genebench-pro-paper-v1-full-pro-extended
GPT-5.4 Pro (Extended)16.3genebench-pro-paper-v1-full-pro-extended
GPT-5.5 Pro (Extended)20.5genebench-pro-paper-v1-full-pro-extended
GPT-5.6 Luna Pro (Extended)23.6genebench-pro-paper-v1-full-pro-extended
GPT-5.6 Terra Pro (Extended)28.5genebench-pro-paper-v1-full-pro-extended
GPT-5.6 Sol Pro (Extended)31.5genebench-pro-paper-v1-full-pro-extended

Eval-level pass rate

genebench-pro-reasoning-enabled · genebench-pro-paper-v1-full-reasoning-enabled

CSV ↓
Accessible data table
ModelValueComparability group
Kimi K2.7 Code2.3genebench-pro-paper-v1-full-reasoning-enabled

Eval-level pass rate

genebench-pro-standard-high · genebench-pro-paper-v1-full-standard-high

CSV ↓
Accessible data table
ModelValueComparability group
Grok 4.31.5genebench-pro-paper-v1-full-standard-high
Gemini 3.1 Pro3.1genebench-pro-paper-v1-full-standard-high
GLM 5.24.6genebench-pro-paper-v1-full-standard-high
Gemini 3.5 Flash8.1genebench-pro-paper-v1-full-standard-high
GPT-5.23.5genebench-pro-paper-v1-full-standard-high
GPT-5.47genebench-pro-paper-v1-full-standard-high
GPT-5.59.3genebench-pro-paper-v1-full-standard-high
GPT-5.6 Luna8genebench-pro-paper-v1-full-standard-high
GPT-5.6 Terra16.2genebench-pro-paper-v1-full-standard-high
GPT-5.6 Sol24.4genebench-pro-paper-v1-full-standard-high

Eval-level pass rate

genebench-pro-standard-low · genebench-pro-paper-v1-full-standard-low

CSV ↓
Accessible data table
ModelValueComparability group
GPT-5.21.1genebench-pro-paper-v1-full-standard-low
GPT-5.43genebench-pro-paper-v1-full-standard-low
GPT-5.52.4genebench-pro-paper-v1-full-standard-low
GPT-5.6 Luna2.3genebench-pro-paper-v1-full-standard-low
GPT-5.6 Terra6.5genebench-pro-paper-v1-full-standard-low
GPT-5.6 Sol14.4genebench-pro-paper-v1-full-standard-low

Eval-level pass rate

genebench-pro-standard-max · genebench-pro-paper-v1-full-standard-max

CSV ↓
Accessible data table
ModelValueComparability group
GPT-5.6 Luna16.5genebench-pro-paper-v1-full-standard-max
GPT-5.6 Terra23.3genebench-pro-paper-v1-full-standard-max
GPT-5.6 Sol28.7genebench-pro-paper-v1-full-standard-max

Eval-level pass rate

genebench-pro-standard-medium · genebench-pro-paper-v1-full-standard-medium

CSV ↓
Accessible data table
ModelValueComparability group
GPT-5.22.4genebench-pro-paper-v1-full-standard-medium
GPT-5.45genebench-pro-paper-v1-full-standard-medium
GPT-5.55.9genebench-pro-paper-v1-full-standard-medium
GPT-5.6 Luna4.7genebench-pro-paper-v1-full-standard-medium
GPT-5.6 Terra13.6genebench-pro-paper-v1-full-standard-medium
GPT-5.6 Sol22.5genebench-pro-paper-v1-full-standard-medium

Eval-level pass rate

genebench-pro-standard-none · genebench-pro-paper-v1-full-standard-none

CSV ↓
Accessible data table
ModelValueComparability group
GPT-5.20.5genebench-pro-paper-v1-full-standard-none
GPT-5.40.9genebench-pro-paper-v1-full-standard-none
GPT-5.50.8genebench-pro-paper-v1-full-standard-none
GPT-5.6 Luna0.8genebench-pro-paper-v1-full-standard-none
GPT-5.6 Terra1genebench-pro-paper-v1-full-standard-none
GPT-5.6 Sol3.7genebench-pro-paper-v1-full-standard-none

Evidence and change history

Source locators remain visible; expand an item to inspect the exact Registry fields it supports.

GeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicine · section: Abstract; Benchmark Scope and Construction; Construction, Validation, and Grading; Methods; Supplementary Tables 1–2 (Definition, 129 total, 10/50/69 release partition, 10 domains and 21 subdomains, 82/47 review strata, environment, grading, attempts, metrics, and results.) · Supports 27 fields

Open source →

  • /name
  • /aliases
  • /summary
  • /kind
  • /organizations
  • /release_date
  • /latest_version
  • /domains
  • /capabilities
  • /modalities
  • /task_formats
  • /task_counts/total
  • /task_counts/basis
  • /task_counts/subsets
  • /coverage_notes
  • /access/level
  • /access/tasks
  • /access/artifacts
  • /access/grader
  • /access/biosafety_notes
  • /resources
  • /implementations
  • /versions/0/task_counts/total
  • /versions/0/task_counts/basis
  • /versions/0/task_counts/subsets
  • /scientific_task_classification/entries/0
  • /scientific_task_classification/entries/1
genebench-pro-public-dataset-resource · repository-path: README.md; LICENSE; problems.csv; manifest.json; reference_grader.py at 9bd2c54a6c0beef041e3504aa7eb65fc77783e18 (Ten public packages, released artifacts, runner contract, commit pin, and the CC-BY-4.0 versus MIT license conflict.) · Supports 9 fields

Open source →

  • /latest_version
  • /access/level
  • /access/tasks
  • /access/artifacts
  • /access/grader
  • /access/license
  • /resources
  • /implementations
  • /versions/0/task_counts/subsets

Unresolved field claims

  • /access/licenseConflicted · high — The pinned Hugging Face README front matter declares CC-BY-4.0, while the same package root LICENSE is the MIT License with OpenAI copyright.
    Evidence: genebench-pro-public-package-evidence

View source-level modification history on GitHub →