Exact model identity
Codex CLI (GPT-5.4)
OpenAI · version status: reported
Similar model names are never merged automatically. Evaluation membership and numeric result rows only reference the exact ID
codex-cli-gpt-5-4.Evaluation settings
| Benchmark | Work | Run / group | Scope |
|---|---|---|---|
| CompBioBench | Agentic systems are adept at solving well-scoped, verifiable problems in computational biology | compbiobench-codex-hardest compbiobench-v1-hardest-codex-xhigh-three-runs | subset · n=17 |
| CompBioBench | Agentic systems are adept at solving well-scoped, verifiable problems in computational biology | compbiobench-creator-full compbiobench-v1-codex-xhigh-three-runs | full · n=100 |
Numeric results
| Benchmark | Work | Run / group | Metric | Value |
|---|---|---|---|---|
| CompBioBench | Agentic systems are adept at solving well-scoped, verifiable problems in computational biology | compbiobench-codex-hardest compbiobench-v1-hardest-codex-xhigh-three-runs | Accuracy — difficulty Levels 4–5 | 59 percent |
| CompBioBench | Agentic systems are adept at solving well-scoped, verifiable problems in computational biology | compbiobench-creator-full compbiobench-v1-codex-xhigh-three-runs | Accuracy | 83.3 percent |
| CompBioBench | Agentic systems are adept at solving well-scoped, verifiable problems in computational biology | compbiobench-creator-full compbiobench-v1-codex-xhigh-three-runs | Wall-clock time per question | 679 seconds |
| CompBioBench | Agentic systems are adept at solving well-scoped, verifiable problems in computational biology | compbiobench-creator-full compbiobench-v1-codex-xhigh-three-runs | Cost per question | 1 USD |