OTIS Mock AIME 2024-2025

Competition mathematics on mock AIME problems written for a training programme rather than scraped from a public archive.

contamination: lowmetric: accuracy

Comparable set

otis-mock-aime-2024-2025/2024-2025/full/0shot/exact-match/inspect

ModelScoreUncertaintyRun byRetrieved
GPT-5.6 Sol100.0%±0.0000third-party2026-08-10
GPT-5.5100.0%±0.0000third-party2026-08-10
Claude Fable 5100.0%±0.0000third-party2026-08-10
GPT-5.6 Terra99.7%±0.0028third-party2026-08-10
Qwen3.8-Max99.4%±0.0039third-party2026-08-10
Claude Opus 598.9%±0.0111third-party2026-08-10
GPT-5.6 Luna98.3%±0.0117third-party2026-08-10
Claude Opus 4.898.3%±0.0141third-party2026-08-10
GPT-5.497.8%±0.0222third-party2026-08-10
Grok 4.597.8%±0.0127third-party2026-08-10
Kimi-K397.2%unknownthird-party2026-08-10
DeepSeek-V4-Pro96.7%±0.0200third-party2026-08-10
Kimi-K2.696.1%±0.0238third-party2026-08-10
Gemini 3.1 Pro (preview)95.6%±0.0310third-party2026-08-10
Qwen3.7-Max95.6%±0.0311third-party2026-08-10
Kimi-K2.7-Code95.6%±0.0311third-party2026-08-10
Gemini 3.5 Flash95.6%±0.0267third-party2026-08-10
Claude Sonnet 594.7%±0.0226third-party2026-08-10
DeepSeek-V4-Flash-073194.4%unknownthird-party2026-08-10
Gemini 3.6 Flash94.2%±0.0309third-party2026-08-10
Grok 4.393.3%±0.0304third-party2026-08-10
GLM-5.192.2%±0.0354third-party2026-08-10
Qwen3.6-27B91.1%±0.0429third-party2026-08-10
GPT-5.4 mini88.9%±0.0474third-party2026-08-10
Qwen3.5-397B-A17B88.9%±0.0474third-party2026-08-10
gpt-oss-120b88.9%unknownthird-party2026-08-10
GPT-5.4 nano87.8%±0.0394third-party2026-08-10
Qwen3.6-35B-A3B86.7%±0.0512third-party2026-08-10
Qwen3-235B-A22B-Instruct-250786.7%±0.0512third-party2026-08-10
GLM-5.286.4%±0.0422third-party2026-08-10
GLM-4.783.3%unknownthird-party2026-08-10
gemma-4-31B-it73.3%±0.0667third-party2026-08-10
Gemini 3.5 Flash Lite71.1%±0.0683third-party2026-08-10
DeepSeek-R1-052866.4%unknownthird-party2026-08-10
gpt-oss-20b53.9%unknownthird-party2026-08-10
gemma-3-27b-it22.2%±0.0540third-party2026-08-10
Llama-4-Maverick-17B-128E-Instruct20.6%unknownthird-party2026-08-10
phi-413.8%unknownthird-party2026-08-10
Qwen2.5-72B-Instruct8.1%unknownthird-party2026-08-10
Llama-4-Scout-17B-16E-Instruct7.8%unknownthird-party2026-08-10
Llama-3.3-70B-Instruct5.1%unknownthird-party2026-08-10
Llama-3.1-70B-Instruct3.6%unknownthird-party2026-08-10
Llama-3.1-8B-Instruct2.5%unknownthird-party2026-08-10
gemma-2-9b-it0.6%unknownthird-party2026-08-10

Epoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. · license CC-BY-4.0 · epoch-ai · https://epoch.ai/benchmarks