Exact model identity
Grok 4.3
xAI · version status: reported
Similar model names are never merged automatically. Evaluation membership and numeric result rows only reference the exact ID
grok-4-3.Evaluation settings
| Benchmark | Work | Run / group | Scope |
|---|---|---|---|
| BioSecBench-Surveillance | BioSecBench-Surveillance repository result snapshot | biosecbench-8d53fd8-pi biosecbench-8d53fd8-pi | full · n=100 |
| GeneBench-Pro | GeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicine | genebench-pro-standard-high genebench-pro-paper-v1-full-standard-high | full · n=129 |
| LifeSciBench | LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences | lifescibench-official-full lifescibench-initial-release-full-official | full · n=750 |
Numeric results
| Benchmark | Work | Run / group | Metric | Value |
|---|---|---|---|---|
| BioSecBench-Surveillance | BioSecBench-Surveillance repository result snapshot | biosecbench-8d53fd8-pi biosecbench-8d53fd8-pi | endpoint pass rate | 16 % |
| GeneBench-Pro | GeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicine | genebench-pro-standard-high genebench-pro-paper-v1-full-standard-high | Eval-level pass rate | 1.5 percent |
| LifeSciBench | LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences | lifescibench-official-full lifescibench-initial-release-full-official | Normalized rubric score | 0.399 proportion |
| LifeSciBench | LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences | lifescibench-official-full lifescibench-initial-release-full-official | Task pass rate | 13 percent |