FrontierMath

Original research-level mathematics problems with automatically verifiable answers. The problem set is held back rather than published.

contamination: lowmetric: accuracy

Comparable set

frontiermath/v1/full/0shot/exact-match/inspect

ModelScoreUncertaintyRun byRetrieved
GPT-5.552.4%±0.0290third-party2026-08-10
GPT-5.450.0%±0.0290third-party2026-08-10
Claude Opus 4.847.2%±0.0294third-party2026-08-10
Gemini 3.5 Flash39.0%±0.0287third-party2026-08-10
Kimi-K2.639.0%±0.0287third-party2026-08-10
Gemini 3.1 Pro (preview)36.9%±0.0280third-party2026-08-10
GLM-5.133.5%±0.0278third-party2026-08-10
GPT-5.4 mini28.3%±0.0207third-party2026-08-10
GPT-5.4 nano25.9%±0.0258third-party2026-08-10
Qwen3-235B-A22B-Instruct-25078.5%±0.0166third-party2026-08-10
GLM-4.72.4%unknownthird-party2026-08-10
Llama-4-Maverick-17B-128E-Instruct0.7%unknownthird-party2026-08-10
Llama-4-Scout-17B-16E-Instruct0.0%unknownthird-party2026-08-10

Epoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. · license CC-BY-4.0 · epoch-ai · https://epoch.ai/benchmarks