OTIS Mock AIME 2024-2025
Competition mathematics on mock AIME problems written for a training programme rather than scraped from a public archive.
contamination: lowmetric: accuracy
Comparable set
otis-mock-aime-2024-2025/2024-2025/full/0shot/exact-match/inspect
| Model | Score | Uncertainty | Run by | Retrieved |
|---|---|---|---|---|
| GPT-5.6 Sol | 100.0% | ±0.0000 | third-party | 2026-08-10 |
| GPT-5.5 | 100.0% | ±0.0000 | third-party | 2026-08-10 |
| Claude Fable 5 | 100.0% | ±0.0000 | third-party | 2026-08-10 |
| GPT-5.6 Terra | 99.7% | ±0.0028 | third-party | 2026-08-10 |
| Qwen3.8-Max | 99.4% | ±0.0039 | third-party | 2026-08-10 |
| Claude Opus 5 | 98.9% | ±0.0111 | third-party | 2026-08-10 |
| GPT-5.6 Luna | 98.3% | ±0.0117 | third-party | 2026-08-10 |
| Claude Opus 4.8 | 98.3% | ±0.0141 | third-party | 2026-08-10 |
| GPT-5.4 | 97.8% | ±0.0222 | third-party | 2026-08-10 |
| Grok 4.5 | 97.8% | ±0.0127 | third-party | 2026-08-10 |
| Kimi-K3 | 97.2% | unknown | third-party | 2026-08-10 |
| DeepSeek-V4-Pro | 96.7% | ±0.0200 | third-party | 2026-08-10 |
| Kimi-K2.6 | 96.1% | ±0.0238 | third-party | 2026-08-10 |
| Gemini 3.1 Pro (preview) | 95.6% | ±0.0310 | third-party | 2026-08-10 |
| Qwen3.7-Max | 95.6% | ±0.0311 | third-party | 2026-08-10 |
| Kimi-K2.7-Code | 95.6% | ±0.0311 | third-party | 2026-08-10 |
| Gemini 3.5 Flash | 95.6% | ±0.0267 | third-party | 2026-08-10 |
| Claude Sonnet 5 | 94.7% | ±0.0226 | third-party | 2026-08-10 |
| DeepSeek-V4-Flash-0731 | 94.4% | unknown | third-party | 2026-08-10 |
| Gemini 3.6 Flash | 94.2% | ±0.0309 | third-party | 2026-08-10 |
| Grok 4.3 | 93.3% | ±0.0304 | third-party | 2026-08-10 |
| GLM-5.1 | 92.2% | ±0.0354 | third-party | 2026-08-10 |
| Qwen3.6-27B | 91.1% | ±0.0429 | third-party | 2026-08-10 |
| GPT-5.4 mini | 88.9% | ±0.0474 | third-party | 2026-08-10 |
| Qwen3.5-397B-A17B | 88.9% | ±0.0474 | third-party | 2026-08-10 |
| gpt-oss-120b | 88.9% | unknown | third-party | 2026-08-10 |
| GPT-5.4 nano | 87.8% | ±0.0394 | third-party | 2026-08-10 |
| Qwen3.6-35B-A3B | 86.7% | ±0.0512 | third-party | 2026-08-10 |
| Qwen3-235B-A22B-Instruct-2507 | 86.7% | ±0.0512 | third-party | 2026-08-10 |
| GLM-5.2 | 86.4% | ±0.0422 | third-party | 2026-08-10 |
| GLM-4.7 | 83.3% | unknown | third-party | 2026-08-10 |
| gemma-4-31B-it | 73.3% | ±0.0667 | third-party | 2026-08-10 |
| Gemini 3.5 Flash Lite | 71.1% | ±0.0683 | third-party | 2026-08-10 |
| DeepSeek-R1-0528 | 66.4% | unknown | third-party | 2026-08-10 |
| gpt-oss-20b | 53.9% | unknown | third-party | 2026-08-10 |
| gemma-3-27b-it | 22.2% | ±0.0540 | third-party | 2026-08-10 |
| Llama-4-Maverick-17B-128E-Instruct | 20.6% | unknown | third-party | 2026-08-10 |
| phi-4 | 13.8% | unknown | third-party | 2026-08-10 |
| Qwen2.5-72B-Instruct | 8.1% | unknown | third-party | 2026-08-10 |
| Llama-4-Scout-17B-16E-Instruct | 7.8% | unknown | third-party | 2026-08-10 |
| Llama-3.3-70B-Instruct | 5.1% | unknown | third-party | 2026-08-10 |
| Llama-3.1-70B-Instruct | 3.6% | unknown | third-party | 2026-08-10 |
| Llama-3.1-8B-Instruct | 2.5% | unknown | third-party | 2026-08-10 |
| gemma-2-9b-it | 0.6% | unknown | third-party | 2026-08-10 |
Epoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. · license CC-BY-4.0 · epoch-ai · https://epoch.ai/benchmarks