Exact model identity

GPT-5.5

OpenAI · version status: reported

Similar model names are never merged automatically. Evaluation membership and numeric result rows only reference the exact ID gpt-5-5.

Evaluation settings

BenchmarkWorkRun / groupScope
BioSecBench-SurveillanceBioSecBench-Surveillance repository result snapshotbiosecbench-8d53fd8-openai-codex
biosecbench-8d53fd8-openai-codex
full · n=100
BioSecBench-SurveillanceBioSecBench-Surveillance repository result snapshotbiosecbench-8d53fd8-pi
biosecbench-8d53fd8-pi
full · n=100
GeneBench-ProGeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicinegenebench-pro-official
genebench-pro-paper-v1-full-standard-xhigh
full · n=129
GeneBench-ProGeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicinegenebench-pro-standard-high
genebench-pro-paper-v1-full-standard-high
full · n=129
GeneBench-ProGeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicinegenebench-pro-standard-low
genebench-pro-paper-v1-full-standard-low
full · n=129
GeneBench-ProGeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicinegenebench-pro-standard-medium
genebench-pro-paper-v1-full-standard-medium
full · n=129
GeneBench-ProGeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicinegenebench-pro-standard-none
genebench-pro-paper-v1-full-standard-none
full · n=129
LifeSciBenchLifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Scienceslifescibench-official-full
lifescibench-initial-release-full-official
full · n=750
SpatialBenchSpatialBench 159-evaluation repository snapshotspatialbench-repo-159-mini-swe-agent
spatialbench-repo-159-mini-swe-agent
full · n=159
SpatialBenchSpatialBench 159-evaluation repository snapshotspatialbench-repo-159-openai-codex
spatialbench-repo-159-openai-codex
full · n=159

Numeric results

BenchmarkWorkRun / groupMetricValue
BioSecBench-SurveillanceBioSecBench-Surveillance repository result snapshotbiosecbench-8d53fd8-openai-codex
biosecbench-8d53fd8-openai-codex
endpoint pass rate50.2 %
BioSecBench-SurveillanceBioSecBench-Surveillance repository result snapshotbiosecbench-8d53fd8-pi
biosecbench-8d53fd8-pi
endpoint pass rate44.8 %
GeneBench-ProGeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicinegenebench-pro-official
genebench-pro-paper-v1-full-standard-xhigh
Eval-level pass rate12 percent
GeneBench-ProGeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicinegenebench-pro-standard-high
genebench-pro-paper-v1-full-standard-high
Eval-level pass rate9.3 percent
GeneBench-ProGeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicinegenebench-pro-standard-low
genebench-pro-paper-v1-full-standard-low
Eval-level pass rate2.4 percent
GeneBench-ProGeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicinegenebench-pro-standard-medium
genebench-pro-paper-v1-full-standard-medium
Eval-level pass rate5.9 percent
GeneBench-ProGeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicinegenebench-pro-standard-none
genebench-pro-paper-v1-full-standard-none
Eval-level pass rate0.8 percent
LifeSciBenchLifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Scienceslifescibench-official-full
lifescibench-initial-release-full-official
Normalized rubric score0.519 proportion
LifeSciBenchLifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Scienceslifescibench-official-full
lifescibench-initial-release-full-official
Task pass rate25.7 percent
SpatialBenchSpatialBench 159-evaluation repository snapshotspatialbench-repo-159-mini-swe-agent
spatialbench-repo-159-mini-swe-agent
Accuracy57.65 percent
SpatialBenchSpatialBench 159-evaluation repository snapshotspatialbench-repo-159-mini-swe-agent
spatialbench-repo-159-mini-swe-agent
Cost1.1207 USD per evaluation
SpatialBenchSpatialBench 159-evaluation repository snapshotspatialbench-repo-159-mini-swe-agent
spatialbench-repo-159-mini-swe-agent
Duration586.66 seconds per evaluation
SpatialBenchSpatialBench 159-evaluation repository snapshotspatialbench-repo-159-openai-codex
spatialbench-repo-159-openai-codex
Accuracy53.67 percent
SpatialBenchSpatialBench 159-evaluation repository snapshotspatialbench-repo-159-openai-codex
spatialbench-repo-159-openai-codex
Cost3.1616 USD per evaluation
SpatialBenchSpatialBench 159-evaluation repository snapshotspatialbench-repo-159-openai-codex
spatialbench-repo-159-openai-codex
Duration382.01 seconds per evaluation