FrontierMath
Original research-level mathematics problems with automatically verifiable answers. The problem set is held back rather than published.
contamination: lowmetric: accuracy
Comparable set
frontiermath/v1/full/0shot/exact-match/inspect
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| GPT-5.5 | 52.4% | ±0.0290 | third-party | 2026-08-10 |
| GPT-5.4 | 50.0% | ±0.0290 | third-party | 2026-08-10 |
| Claude Opus 4.8 | 47.2% | ±0.0294 | third-party | 2026-08-10 |
| Gemini 3.5 Flash | 39.0% | ±0.0287 | third-party | 2026-08-10 |
| Kimi-K2.6 | 39.0% | ±0.0287 | third-party | 2026-08-10 |
| Gemini 3.1 Pro (preview) | 36.9% | ±0.0280 | third-party | 2026-08-10 |
| GLM-5.1 | 33.5% | ±0.0278 | third-party | 2026-08-10 |
| GPT-5.4 mini | 28.3% | ±0.0207 | third-party | 2026-08-10 |
| GPT-5.4 nano | 25.9% | ±0.0258 | third-party | 2026-08-10 |
| Qwen3-235B-A22B-Instruct-2507 | 8.5% | ±0.0166 | third-party | 2026-08-10 |
| GLM-4.7 | 2.4% | unknown | third-party | 2026-08-10 |
| Llama-4-Maverick-17B-128E-Instruct | 0.7% | unknown | third-party | 2026-08-10 |
| Llama-4-Scout-17B-16E-Instruct | 0.0% | unknown | third-party | 2026-08-10 |
Epoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. · license CC-BY-4.0 · epoch-ai · https://epoch.ai/benchmarks