preprint · benchmark creator

GeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicine

OpenAI · 2026-06-30

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

No BenchmarkUse relation is normalized for this legacy work yet. Existing EvaluationRuns remain available below.

Normalized evaluation runs

GeneBench-Pro13 runs

Open benchmark record →

genebench-pro-paper-v1-full-claude-highvpaper-v1

Evaluated models / systems: Claude Opus 4.8

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortprovider-reported high reasoning setting
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats5
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
Claude Opus 4.8Eval-level pass rate9 percent
Supplementary Table 1 configuration ClaudeOpus4.8 (high); nominal attempts per problem 5, valid-attempt mean 5.0 and range 4–5; average tokens not reported; problem regimes 0%=78.3%, 0–10%=0.0%, 10–50%=14.7%, ≥50%=7.0%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-claude-lowvpaper-v1

Evaluated models / systems: Claude Opus 4.8

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortprovider-reported low reasoning setting
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats5
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
Claude Opus 4.8Eval-level pass rate4.3 percent
Supplementary Table 1 configuration ClaudeOpus4.8 (low); nominal attempts per problem 5, valid-attempt mean 5.0 and range 3–5; average tokens not reported; problem regimes 0%=88.4%, 0–10%=0.0%, 10–50%=10.1%, ≥50%=1.6%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-claude-maxvpaper-v1

Evaluated models / systems: Claude Opus 4.8

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortprovider-reported max reasoning setting
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats5
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
Claude Opus 4.8Eval-level pass rate16 percent
Supplementary Table 1 configuration ClaudeOpus4.8 (max); nominal attempts per problem 5, valid-attempt mean 5.0 and range 5–5; average tokens not reported; problem regimes 0%=67.4%, 0–10%=0.0%, 10–50%=18.6%, ≥50%=14.0%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-claude-mediumvpaper-v1

Evaluated models / systems: Claude Opus 4.8

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortprovider-reported medium reasoning setting
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats5
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
Claude Opus 4.8Eval-level pass rate4.3 percent
Supplementary Table 1 configuration ClaudeOpus4.8 (medium); nominal attempts per problem 5, valid-attempt mean 5.0 and range 5–5; average tokens not reported; problem regimes 0%=87.6%, 0–10%=0.0%, 10–50%=10.1%, ≥50%=2.3%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-claude-xhighvpaper-v1

Evaluated models / systems: Claude Opus 4.8

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortprovider-reported xhigh reasoning setting
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats5
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
Claude Opus 4.8Eval-level pass rate10.1 percent
Supplementary Table 1 configuration ClaudeOpus4.8 (xhigh); nominal attempts per problem 5, valid-attempt mean 5.0 and range 4–5; average tokens not reported; problem regimes 0%=75.2%, 0–10%=0.0%, 10–50%=17.1%, ≥50%=7.8%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-standard-xhighvpaper-v1

Evaluated models / systems: DeepSeek V4 Flash, DeepSeek V4 Pro, GLM 5.1, GPT-5.2, GPT-5.4, GPT-5.5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, Kimi K2.6, MiMo V2.5, MiMo V2.5 Pro, MiniMax M2.7, MiniMax M3, Qwen 3.7 Max, Qwen 3.7 Plus, Tencent HY 3 Preview

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortprovider-reported xhigh reasoning setting
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats10
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
MiniMax M2.7Eval-level pass rate0.6 percent
Supplementary Table 1 configuration MiniMaxM2.7 (xhigh); nominal attempts per problem 10, valid-attempt mean 9.9 and range 8–10; average tokens not reported; problem regimes 0%=96.1%, 0–10%=2.3%, 10–50%=1.6%, ≥50%=0.0%.
129
Tencent HY 3 PreviewEval-level pass rate0.9 percent
Supplementary Table 1 configuration TencentHY3Preview (xhigh); nominal attempts per problem 10, valid-attempt mean 9.8 and range 8–10; average tokens not reported; problem regimes 0%=92.2%, 0–10%=6.2%, 10–50%=1.6%, ≥50%=0.0%.
129
MiniMax M3Eval-level pass rate0.9 percent
Supplementary Table 1 configuration MiniMaxM3 (xhigh); nominal attempts per problem 10, valid-attempt mean 9.6 and range 7–10; average tokens not reported; problem regimes 0%=96.1%, 0–10%=1.6%, 10–50%=2.3%, ≥50%=0.0%.
129
MiMo V2.5Eval-level pass rate1.2 percent
Supplementary Table 1 configuration MiMoV2.5 (xhigh); nominal attempts per problem 10, valid-attempt mean 9.9 and range 8–10; average tokens not reported; problem regimes 0%=91.5%, 0–10%=6.2%, 10–50%=2.3%, ≥50%=0.0%.
129
GLM 5.1Eval-level pass rate1.2 percent
Supplementary Table 1 configuration GLM5.1 (xhigh); nominal attempts per problem 10, valid-attempt mean 9.7 and range 8–10; average tokens not reported; problem regimes 0%=95.3%, 0–10%=2.3%, 10–50%=1.6%, ≥50%=0.8%.
129
MiMo V2.5 ProEval-level pass rate2 percent
Supplementary Table 1 configuration MiMoV2.5Pro (xhigh); nominal attempts per problem 10, valid-attempt mean 9.9 and range 8–10; average tokens not reported; problem regimes 0%=86.8%, 0–10%=9.3%, 10–50%=3.1%, ≥50%=0.8%.
129
Qwen 3.7 PlusEval-level pass rate2.3 percent
Supplementary Table 1 configuration Qwen3.7Plus (xhigh); nominal attempts per problem 10, valid-attempt mean 9.7 and range 6–10; average tokens not reported; problem regimes 0%=86.8%, 0–10%=7.0%, 10–50%=5.4%, ≥50%=0.8%.
129
DeepSeek V4 FlashEval-level pass rate2.4 percent
Supplementary Table 1 configuration DeepSeekV4Flash (xhigh); nominal attempts per problem 10, valid-attempt mean 9.9 and range 9–10; average tokens not reported; problem regimes 0%=87.6%, 0–10%=6.2%, 10–50%=4.7%, ≥50%=1.6%.
129
DeepSeek V4 ProEval-level pass rate2.4 percent
Supplementary Table 1 configuration DeepSeekV4Pro (xhigh); nominal attempts per problem 10, valid-attempt mean 9.6 and range 8–10; average tokens not reported; problem regimes 0%=84.5%, 0–10%=4.7%, 10–50%=10.9%, ≥50%=0.0%.
129
Qwen 3.7 MaxEval-level pass rate4 percent
Supplementary Table 1 configuration Qwen3.7Max (xhigh); nominal attempts per problem 10, valid-attempt mean 9.8 and range 9–10; average tokens not reported; problem regimes 0%=79.8%, 0–10%=7.8%, 10–50%=11.6%, ≥50%=0.8%.
129
Kimi K2.6Eval-level pass rate4.4 percent
Supplementary Table 1 configuration KimiK2.6 (xhigh); nominal attempts per problem 10, valid-attempt mean 9.9 and range 7–10; average tokens not reported; problem regimes 0%=84.5%, 0–10%=6.2%, 10–50%=6.2%, ≥50%=3.1%.
129
GPT-5.2Eval-level pass rate4.9 percent
Supplementary Table 1 configuration GPT-5.2 (xhigh); nominal attempts per problem 10, valid-attempt mean 9.9 and range 9–10; average tokens 52.0k; problem regimes 0%=77.5%, 0–10%=10.1%, 10–50%=10.9%, ≥50%=1.6%.
129
GPT-5.4Eval-level pass rate8.9 percent
Supplementary Table 1 configuration GPT-5.4 (xhigh); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 44.7k; problem regimes 0%=67.4%, 0–10%=9.3%, 10–50%=18.6%, ≥50%=4.7%.
129
GPT-5.5Eval-level pass rate12 percent
Supplementary Table 1 configuration GPT-5.5 (xhigh); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 28.7k; problem regimes 0%=64.3%, 0–10%=8.5%, 10–50%=18.6%, ≥50%=8.5%.
129
GPT-5.6 LunaEval-level pass rate10.8 percent
Supplementary Table 1 configuration GPT-5.6Luna (xhigh); nominal attempts per problem 10, valid-attempt mean 9.8 and range 7–10; average tokens 53.1k; problem regimes 0%=70.5%, 0–10%=8.5%, 10–50%=10.1%, ≥50%=10.9%.
129
GPT-5.6 TerraEval-level pass rate18.8 percent
Supplementary Table 1 configuration GPT-5.6Terra (xhigh); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 31.1k; problem regimes 0%=56.6%, 0–10%=7.8%, 10–50%=19.4%, ≥50%=16.3%.
129
GPT-5.6 SolEval-level pass rate26.8 percent
Supplementary Table 1 configuration GPT-5.6Sol (xhigh); nominal attempts per problem 10, valid-attempt mean 9.9 and range 9–10; average tokens 25.7k; problem regimes 0%=50.4%, 0–10%=6.2%, 10–50%=14.0%, ≥50%=29.5%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-pro-extendedvpaper-v1

Evaluated models / systems: GPT-5.2 Pro (Extended), GPT-5.4 Pro (Extended), GPT-5.5 Pro (Extended), GPT-5.6 Luna Pro (Extended), GPT-5.6 Sol Pro (Extended), GPT-5.6 Terra Pro (Extended)

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortGPT Pro (Extended)
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats5
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
GPT-5.2 Pro (Extended)Eval-level pass rate8.5 percent
Supplementary Table 1 configuration GPT-5.2Pro (Extended); nominal attempts per problem 5, valid-attempt mean 5.0 and range 4–5; average tokens not reported; problem regimes 0%=79.1%, 0–10%=0.0%, 10–50%=14.0%, ≥50%=7.0%.
129
GPT-5.4 Pro (Extended)Eval-level pass rate16.3 percent
Supplementary Table 1 configuration GPT-5.4Pro (Extended); nominal attempts per problem 5, valid-attempt mean 5.0 and range 5–5; average tokens not reported; problem regimes 0%=72.1%, 0–10%=0.0%, 10–50%=12.4%, ≥50%=15.5%.
129
GPT-5.5 Pro (Extended)Eval-level pass rate20.5 percent
Supplementary Table 1 configuration GPT-5.5Pro (Extended); nominal attempts per problem 5, valid-attempt mean 5.0 and range 5–5; average tokens not reported; problem regimes 0%=66.7%, 0–10%=0.0%, 10–50%=14.0%, ≥50%=19.4%.
129
GPT-5.6 Luna Pro (Extended)Eval-level pass rate23.6 percent
Supplementary Table 1 configuration GPT-5.6LunaPro (Extended); nominal attempts per problem 5, valid-attempt mean 5.0 and range 5–5; average tokens not reported; problem regimes 0%=65.9%, 0–10%=0.0%, 10–50%=10.9%, ≥50%=23.3%.
129
GPT-5.6 Terra Pro (Extended)Eval-level pass rate28.5 percent
Supplementary Table 1 configuration GPT-5.6TerraPro (Extended); nominal attempts per problem 5, valid-attempt mean 5.0 and range 5–5; average tokens not reported; problem regimes 0%=59.7%, 0–10%=0.0%, 10–50%=11.6%, ≥50%=28.7%.
129
GPT-5.6 Sol Pro (Extended)Eval-level pass rate31.5 percent
Supplementary Table 1 configuration GPT-5.6SolPro (Extended); nominal attempts per problem 5, valid-attempt mean 5.0 and range 5–5; average tokens not reported; problem regimes 0%=55.8%, 0–10%=0.0%, 10–50%=14.0%, ≥50%=30.2%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-reasoning-enabledvpaper-v1

Evaluated models / systems: Kimi K2.7 Code

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortreasoning_enabled
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats10
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
Kimi K2.7 CodeEval-level pass rate2.3 percent
Supplementary Table 1 configuration KimiK2.7Code (reasoning_enabled); nominal attempts per problem 10, valid-attempt mean 9.9 and range 9–10; average tokens not reported; problem regimes 0%=84.5%, 0–10%=9.3%, 10–50%=6.2%, ≥50%=0.0%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-standard-highvpaper-v1

Evaluated models / systems: Gemini 3.1 Pro, Gemini 3.5 Flash, GLM 5.2, GPT-5.2, GPT-5.4, GPT-5.5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, Grok 4.3

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortprovider-reported high reasoning setting
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats10
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
Grok 4.3Eval-level pass rate1.5 percent
Supplementary Table 1 configuration Grok4.3 (high); nominal attempts per problem 10, valid-attempt mean 10.0 and range 10–10; average tokens not reported; problem regimes 0%=92.2%, 0–10%=4.7%, 10–50%=2.3%, ≥50%=0.8%.
129
Gemini 3.1 ProEval-level pass rate3.1 percent
Supplementary Table 1 configuration Gemini3.1Pro (high); nominal attempts per problem 10, valid-attempt mean 10.0 and range 10–10; average tokens not reported; problem regimes 0%=81.4%, 0–10%=13.2%, 10–50%=4.7%, ≥50%=0.8%.
129
GLM 5.2Eval-level pass rate4.6 percent
Supplementary Table 1 configuration GLM5.2 (high); nominal attempts per problem 10, valid-attempt mean 9.8 and range 8–10; average tokens not reported; problem regimes 0%=77.5%, 0–10%=10.1%, 10–50%=10.9%, ≥50%=1.6%.
129
Gemini 3.5 FlashEval-level pass rate8.1 percent
Supplementary Table 1 configuration Gemini3.5Flash (high); nominal attempts per problem 10, valid-attempt mean 10.0 and range 10–10; average tokens not reported; problem regimes 0%=70.5%, 0–10%=14.7%, 10–50%=9.3%, ≥50%=5.4%.
129
GPT-5.2Eval-level pass rate3.5 percent
Supplementary Table 1 configuration GPT-5.2 (high); nominal attempts per problem 10, valid-attempt mean 9.9 and range 9–10; average tokens 23.2k; problem regimes 0%=83.7%, 0–10%=7.8%, 10–50%=7.0%, ≥50%=1.6%.
129
GPT-5.4Eval-level pass rate7 percent
Supplementary Table 1 configuration GPT-5.4 (high); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 27.3k; problem regimes 0%=70.5%, 0–10%=11.6%, 10–50%=13.2%, ≥50%=4.7%.
129
GPT-5.5Eval-level pass rate9.3 percent
Supplementary Table 1 configuration GPT-5.5 (high); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 19.8k; problem regimes 0%=70.5%, 0–10%=11.6%, 10–50%=10.9%, ≥50%=7.0%.
129
GPT-5.6 LunaEval-level pass rate8 percent
Supplementary Table 1 configuration GPT-5.6Luna (high); nominal attempts per problem 10, valid-attempt mean 9.8 and range 7–10; average tokens 32.3k; problem regimes 0%=76.7%, 0–10%=8.5%, 10–50%=7.8%, ≥50%=7.0%.
129
GPT-5.6 TerraEval-level pass rate16.2 percent
Supplementary Table 1 configuration GPT-5.6Terra (high); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 22.2k; problem regimes 0%=59.7%, 0–10%=12.4%, 10–50%=11.6%, ≥50%=16.3%.
129
GPT-5.6 SolEval-level pass rate24.4 percent
Supplementary Table 1 configuration GPT-5.6Sol (high); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 19.5k; problem regimes 0%=51.9%, 0–10%=7.8%, 10–50%=17.8%, ≥50%=22.5%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-standard-lowvpaper-v1

Evaluated models / systems: GPT-5.2, GPT-5.4, GPT-5.5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortprovider-reported low reasoning setting
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats10
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
GPT-5.2Eval-level pass rate1.1 percent
Supplementary Table 1 configuration GPT-5.2 (low); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 7.0k; problem regimes 0%=91.5%, 0–10%=6.2%, 10–50%=2.3%, ≥50%=0.0%.
129
GPT-5.4Eval-level pass rate3 percent
Supplementary Table 1 configuration GPT-5.4 (low); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 8.2k; problem regimes 0%=85.3%, 0–10%=7.0%, 10–50%=7.0%, ≥50%=0.8%.
129
GPT-5.5Eval-level pass rate2.4 percent
Supplementary Table 1 configuration GPT-5.5 (low); nominal attempts per problem 10, valid-attempt mean 9.9 and range 8–10; average tokens 2.8k; problem regimes 0%=87.6%, 0–10%=6.2%, 10–50%=4.7%, ≥50%=1.6%.
129
GPT-5.6 LunaEval-level pass rate2.3 percent
Supplementary Table 1 configuration GPT-5.6Luna (low); nominal attempts per problem 10, valid-attempt mean 10.0 and range 10–10; average tokens 3.6k; problem regimes 0%=89.1%, 0–10%=5.4%, 10–50%=4.7%, ≥50%=0.8%.
129
GPT-5.6 TerraEval-level pass rate6.5 percent
Supplementary Table 1 configuration GPT-5.6Terra (low); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 5.5k; problem regimes 0%=75.2%, 0–10%=10.1%, 10–50%=10.9%, ≥50%=3.9%.
129
GPT-5.6 SolEval-level pass rate14.4 percent
Supplementary Table 1 configuration GPT-5.6Sol (low); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 5.6k; problem regimes 0%=58.9%, 0–10%=12.4%, 10–50%=17.1%, ≥50%=11.6%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-standard-maxvpaper-v1

Evaluated models / systems: GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortprovider-reported max reasoning setting
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats10
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
GPT-5.6 LunaEval-level pass rate16.5 percent
Supplementary Table 1 configuration GPT-5.6Luna (max); nominal attempts per problem 10, valid-attempt mean 9.6 and range 2–10; average tokens 118.2k; problem regimes 0%=64.3%, 0–10%=4.7%, 10–50%=16.3%, ≥50%=14.7%.
129
GPT-5.6 TerraEval-level pass rate23.3 percent
Supplementary Table 1 configuration GPT-5.6Terra (max); nominal attempts per problem 10, valid-attempt mean 10.0 and range 10–10; average tokens 54.3k; problem regimes 0%=49.6%, 0–10%=9.3%, 10–50%=19.4%, ≥50%=21.7%.
129
GPT-5.6 SolEval-level pass rate28.7 percent
Supplementary Table 1 configuration GPT-5.6Sol (max); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 33.2k; problem regimes 0%=45.7%, 0–10%=10.1%, 10–50%=14.0%, ≥50%=30.2%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-standard-mediumvpaper-v1

Evaluated models / systems: GPT-5.2, GPT-5.4, GPT-5.5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortprovider-reported medium reasoning setting
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats10
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
GPT-5.2Eval-level pass rate2.4 percent
Supplementary Table 1 configuration GPT-5.2 (medium); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 18.1k; problem regimes 0%=86.0%, 0–10%=10.1%, 10–50%=3.1%, ≥50%=0.8%.
129
GPT-5.4Eval-level pass rate5 percent
Supplementary Table 1 configuration GPT-5.4 (medium); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 18.6k; problem regimes 0%=74.4%, 0–10%=17.1%, 10–50%=7.0%, ≥50%=1.6%.
129
GPT-5.5Eval-level pass rate5.9 percent
Supplementary Table 1 configuration GPT-5.5 (medium); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 10.6k; problem regimes 0%=79.1%, 0–10%=7.8%, 10–50%=9.3%, ≥50%=3.9%.
129
GPT-5.6 LunaEval-level pass rate4.7 percent
Supplementary Table 1 configuration GPT-5.6Luna (medium); nominal attempts per problem 10, valid-attempt mean 9.9 and range 8–10; average tokens 15.6k; problem regimes 0%=83.7%, 0–10%=7.0%, 10–50%=4.7%, ≥50%=4.7%.
129
GPT-5.6 TerraEval-level pass rate13.6 percent
Supplementary Table 1 configuration GPT-5.6Terra (medium); nominal attempts per problem 10, valid-attempt mean 10.0 and range 10–10; average tokens 15.9k; problem regimes 0%=65.9%, 0–10%=8.5%, 10–50%=12.4%, ≥50%=13.2%.
129
GPT-5.6 SolEval-level pass rate22.5 percent
Supplementary Table 1 configuration GPT-5.6Sol (medium); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 14.4k; problem regimes 0%=53.5%, 0–10%=7.8%, 10–50%=16.3%, ≥50%=22.5%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results
genebench-pro-paper-v1-full-standard-nonevpaper-v1

Evaluated models / systems: GPT-5.2, GPT-5.4, GPT-5.5, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra

Scopefull · n=129
ShotsNot reported
Turnsmulti-turn agent-container trajectory
System prompt publicNo
Reasoning / effortprovider-reported none reasoning setting
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolsPython/R scientific stacks, PLINK 2.0, bedtools/tabix, pysam/cyvcf2, scanpy/anndata, DESeq2/edgeR/limma
Token budgetNot reported
Time / cost budgetNo additional uniform wall-clock budget imposed by the harness
TemperatureNot reported
SeedNot reported
Repeats10
Graderproblem-specific Python binary checker requiring every graded field to satisfy exact-match or absolute-tolerance constraints · human review: no
Statisticsunweighted mean of 129 per-problem pass rates with a 95% hierarchical bootstrap CI from 20,000 resamples of problems and repeated runs
Contaminationconstructively simulated suite with disjoint 10 public, 50 Artificial Analysis, and 69 internal-holdout release strata
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Eval-level pass rateabsolutepercentunweighted mean across 129 problem-level pass rates; each problem-level rate is passing valid attempts divided by valid attemptsA problem attempt passes only when all graded fields satisfy their problem-specific exact-match rules or absolute numeric tolerances.

Results

ModelMetricValuen
GPT-5.2Eval-level pass rate0.5 percent
Supplementary Table 1 configuration GPT-5.2 (none); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 1.1k; problem regimes 0%=96.9%, 0–10%=2.3%, 10–50%=0.8%, ≥50%=0.0%.
129
GPT-5.4Eval-level pass rate0.9 percent
Supplementary Table 1 configuration GPT-5.4 (none); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 2.1k; problem regimes 0%=94.6%, 0–10%=3.9%, 10–50%=1.6%, ≥50%=0.0%.
129
GPT-5.5Eval-level pass rate0.8 percent
Supplementary Table 1 configuration GPT-5.5 (none); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 1.1k; problem regimes 0%=96.1%, 0–10%=2.3%, 10–50%=1.6%, ≥50%=0.0%.
129
GPT-5.6 LunaEval-level pass rate0.8 percent
Supplementary Table 1 configuration GPT-5.6Luna (none); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 975; problem regimes 0%=93.8%, 0–10%=4.7%, 10–50%=1.6%, ≥50%=0.0%.
129
GPT-5.6 TerraEval-level pass rate1 percent
Supplementary Table 1 configuration GPT-5.6Terra (none); nominal attempts per problem 10, valid-attempt mean 10.0 and range 10–10; average tokens 930; problem regimes 0%=95.3%, 0–10%=3.1%, 10–50%=0.8%, ≥50%=0.8%.
129
GPT-5.6 SolEval-level pass rate3.7 percent
Supplementary Table 1 configuration GPT-5.6Sol (none); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 1.4k; problem regimes 0%=82.2%, 0–10%=9.3%, 10–50%=7.0%, ≥50%=1.6%.
129

Evidence

  • section: Methods — Evaluation and grading, printed page 15 (Full scope, attempts, invalid-run exclusion, Docker environment, installed tools, no internet, response schema, binary grader, and no uniform harness wall-clock budget.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Supplementary Table 1, printed page 21 (Exact configuration labels, pass rates, 95% CI bounds, average tokens where available, valid-attempt summaries, and problem-regime shares.) — supports /results