| MiniMax M2.7 | Eval-level pass rate | 0.6 percent Supplementary Table 1 configuration MiniMaxM2.7 (xhigh); nominal attempts per problem 10, valid-attempt mean 9.9 and range 8–10; average tokens not reported; problem regimes 0%=96.1%, 0–10%=2.3%, 10–50%=1.6%, ≥50%=0.0%. | 129 |
| Tencent HY 3 Preview | Eval-level pass rate | 0.9 percent Supplementary Table 1 configuration TencentHY3Preview (xhigh); nominal attempts per problem 10, valid-attempt mean 9.8 and range 8–10; average tokens not reported; problem regimes 0%=92.2%, 0–10%=6.2%, 10–50%=1.6%, ≥50%=0.0%. | 129 |
| MiniMax M3 | Eval-level pass rate | 0.9 percent Supplementary Table 1 configuration MiniMaxM3 (xhigh); nominal attempts per problem 10, valid-attempt mean 9.6 and range 7–10; average tokens not reported; problem regimes 0%=96.1%, 0–10%=1.6%, 10–50%=2.3%, ≥50%=0.0%. | 129 |
| MiMo V2.5 | Eval-level pass rate | 1.2 percent Supplementary Table 1 configuration MiMoV2.5 (xhigh); nominal attempts per problem 10, valid-attempt mean 9.9 and range 8–10; average tokens not reported; problem regimes 0%=91.5%, 0–10%=6.2%, 10–50%=2.3%, ≥50%=0.0%. | 129 |
| GLM 5.1 | Eval-level pass rate | 1.2 percent Supplementary Table 1 configuration GLM5.1 (xhigh); nominal attempts per problem 10, valid-attempt mean 9.7 and range 8–10; average tokens not reported; problem regimes 0%=95.3%, 0–10%=2.3%, 10–50%=1.6%, ≥50%=0.8%. | 129 |
| MiMo V2.5 Pro | Eval-level pass rate | 2 percent Supplementary Table 1 configuration MiMoV2.5Pro (xhigh); nominal attempts per problem 10, valid-attempt mean 9.9 and range 8–10; average tokens not reported; problem regimes 0%=86.8%, 0–10%=9.3%, 10–50%=3.1%, ≥50%=0.8%. | 129 |
| Qwen 3.7 Plus | Eval-level pass rate | 2.3 percent Supplementary Table 1 configuration Qwen3.7Plus (xhigh); nominal attempts per problem 10, valid-attempt mean 9.7 and range 6–10; average tokens not reported; problem regimes 0%=86.8%, 0–10%=7.0%, 10–50%=5.4%, ≥50%=0.8%. | 129 |
| DeepSeek V4 Flash | Eval-level pass rate | 2.4 percent Supplementary Table 1 configuration DeepSeekV4Flash (xhigh); nominal attempts per problem 10, valid-attempt mean 9.9 and range 9–10; average tokens not reported; problem regimes 0%=87.6%, 0–10%=6.2%, 10–50%=4.7%, ≥50%=1.6%. | 129 |
| DeepSeek V4 Pro | Eval-level pass rate | 2.4 percent Supplementary Table 1 configuration DeepSeekV4Pro (xhigh); nominal attempts per problem 10, valid-attempt mean 9.6 and range 8–10; average tokens not reported; problem regimes 0%=84.5%, 0–10%=4.7%, 10–50%=10.9%, ≥50%=0.0%. | 129 |
| Qwen 3.7 Max | Eval-level pass rate | 4 percent Supplementary Table 1 configuration Qwen3.7Max (xhigh); nominal attempts per problem 10, valid-attempt mean 9.8 and range 9–10; average tokens not reported; problem regimes 0%=79.8%, 0–10%=7.8%, 10–50%=11.6%, ≥50%=0.8%. | 129 |
| Kimi K2.6 | Eval-level pass rate | 4.4 percent Supplementary Table 1 configuration KimiK2.6 (xhigh); nominal attempts per problem 10, valid-attempt mean 9.9 and range 7–10; average tokens not reported; problem regimes 0%=84.5%, 0–10%=6.2%, 10–50%=6.2%, ≥50%=3.1%. | 129 |
| GPT-5.2 | Eval-level pass rate | 4.9 percent Supplementary Table 1 configuration GPT-5.2 (xhigh); nominal attempts per problem 10, valid-attempt mean 9.9 and range 9–10; average tokens 52.0k; problem regimes 0%=77.5%, 0–10%=10.1%, 10–50%=10.9%, ≥50%=1.6%. | 129 |
| GPT-5.4 | Eval-level pass rate | 8.9 percent Supplementary Table 1 configuration GPT-5.4 (xhigh); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 44.7k; problem regimes 0%=67.4%, 0–10%=9.3%, 10–50%=18.6%, ≥50%=4.7%. | 129 |
| GPT-5.5 | Eval-level pass rate | 12 percent Supplementary Table 1 configuration GPT-5.5 (xhigh); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 28.7k; problem regimes 0%=64.3%, 0–10%=8.5%, 10–50%=18.6%, ≥50%=8.5%. | 129 |
| GPT-5.6 Luna | Eval-level pass rate | 10.8 percent Supplementary Table 1 configuration GPT-5.6Luna (xhigh); nominal attempts per problem 10, valid-attempt mean 9.8 and range 7–10; average tokens 53.1k; problem regimes 0%=70.5%, 0–10%=8.5%, 10–50%=10.1%, ≥50%=10.9%. | 129 |
| GPT-5.6 Terra | Eval-level pass rate | 18.8 percent Supplementary Table 1 configuration GPT-5.6Terra (xhigh); nominal attempts per problem 10, valid-attempt mean 10.0 and range 9–10; average tokens 31.1k; problem regimes 0%=56.6%, 0–10%=7.8%, 10–50%=19.4%, ≥50%=16.3%. | 129 |
| GPT-5.6 Sol | Eval-level pass rate | 26.8 percent Supplementary Table 1 configuration GPT-5.6Sol (xhigh); nominal attempts per problem 10, valid-attempt mean 9.9 and range 9–10; average tokens 25.7k; problem regimes 0%=50.4%, 0–10%=6.2%, 10–50%=14.0%, ≥50%=29.5%. | 129 |