Exact model identity

GPT-5.2

OpenAI · version status: reported

Similar model names are never merged automatically. Evaluation membership and numeric result rows only reference the exact ID gpt-5-2.

Evaluation settings

Numeric results

BenchmarkWorkRun / groupMetricValue
GeneBench-ProGeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicinegenebench-pro-official
genebench-pro-paper-v1-full-standard-xhigh
Eval-level pass rate4.9 percent
GeneBench-ProGeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicinegenebench-pro-standard-high
genebench-pro-paper-v1-full-standard-high
Eval-level pass rate3.5 percent
GeneBench-ProGeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicinegenebench-pro-standard-low
genebench-pro-paper-v1-full-standard-low
Eval-level pass rate1.1 percent
GeneBench-ProGeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicinegenebench-pro-standard-medium
genebench-pro-paper-v1-full-standard-medium
Eval-level pass rate2.4 percent
GeneBench-ProGeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicinegenebench-pro-standard-none
genebench-pro-paper-v1-full-standard-none
Eval-level pass rate0.5 percent
SpatialBenchSpatialBench: Can Agents Analyze Real-World Spatial Biology Data?spatialbench-paper-v2-base
spatialbench-paper-v2-base
Accuracy34.02 percent
SpatialBenchSpatialBench: Can Agents Analyze Real-World Spatial Biology Data?spatialbench-paper-v2-base
spatialbench-paper-v2-base
Steps2.1 steps per evaluation
SpatialBenchSpatialBench: Can Agents Analyze Real-World Spatial Biology Data?spatialbench-paper-v2-base
spatialbench-paper-v2-base
Latency89.2 seconds per evaluation
SpatialBenchSpatialBench: Can Agents Analyze Real-World Spatial Biology Data?spatialbench-paper-v2-base
spatialbench-paper-v2-base
Cost0.037 USD per evaluation
SpatialBenchSpatialBench 159-evaluation repository snapshotspatialbench-repo-159-mini-swe-agent
spatialbench-repo-159-mini-swe-agent
Accuracy50.1 percent
SpatialBenchSpatialBench 159-evaluation repository snapshotspatialbench-repo-159-mini-swe-agent
spatialbench-repo-159-mini-swe-agent
Cost0.6024 USD per evaluation
SpatialBenchSpatialBench 159-evaluation repository snapshotspatialbench-repo-159-mini-swe-agent
spatialbench-repo-159-mini-swe-agent
Duration931.76 seconds per evaluation