FrontierMath Tier 4
The hardest tier of FrontierMath: problems pitched at research mathematics. Near-floor scores separated by one or two problems are noise.
contamination: lowmetric: accuracy
Comparable set
frontiermath-tier-4/v1/tier-4/0shot/exact-match/inspect
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| GPT-5.5 | 39.6% | ±0.0710 | third-party | 2026-08-10 |
| GPT-5.4 | 37.5% | ±0.0700 | third-party | 2026-08-10 |
| Claude Opus 4.8 | 31.3% | ±0.0676 | third-party | 2026-08-10 |
| Gemini 3.1 Pro (preview) | 16.7% | ±0.0540 | third-party | 2026-08-10 |
| Gemini 3.5 Flash | 14.6% | ±0.0515 | third-party | 2026-08-10 |
| Kimi-K2.6 | 14.6% | ±0.0515 | third-party | 2026-08-10 |
| GLM-5.1 | 12.5% | ±0.0482 | third-party | 2026-08-10 |
| GPT-5.4 nano | 6.3% | ±0.0353 | third-party | 2026-08-10 |
| GPT-5.4 mini | 2.1% | ±0.0208 | third-party | 2026-08-10 |
| GLM-4.7 | 0.0% | unknown | third-party | 2026-08-10 |
| Qwen3-235B-A22B-Instruct-2507 | 0.0% | ±0.0000 | third-party | 2026-08-10 |
Epoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. · license CC-BY-4.0 · epoch-ai · https://epoch.ai/benchmarks