GPQA Diamond

Graduate-Level Google-Proof Q&A (Diamond split). Hard STEM questions designed to be difficult for non-experts with web search.

contamination: lowmetric: accuracy

Comparable set

gpqa-diamond/v1/diamond/0shot/exact-match/inspect

ModelScoreUncertaintyRun byRetrieved
GPT-5.494.6%±0.0160third-party2026-08-10
Gemini 3.1 Pro (preview)94.4%±0.0163third-party2026-08-10
Gemini 3.6 Flash94.1%±0.0140third-party2026-08-10
GPT-5.594.0%±0.0155third-party2026-08-10
Claude Opus 593.9%±0.0148third-party2026-08-10
GPT-5.6 Sol93.5%±0.0157third-party2026-08-10
Grok 4.593.4%±0.0143third-party2026-08-10
GPT-5.6 Terra93.3%±0.0154third-party2026-08-10
Kimi-K393.1%unknownthird-party2026-08-10
Gemini 3.5 Flash92.8%±0.0164third-party2026-08-10
Qwen3.8-Max92.7%±0.0169third-party2026-08-10
GLM-5.291.9%±0.0161third-party2026-08-10
GPT-5.6 Luna91.6%±0.0173third-party2026-08-10
Claude Opus 4.891.0%±0.0192third-party2026-08-10
DeepSeek-V4-Flash-073191.0%unknownthird-party2026-08-10
Qwen3.7-Max90.9%±0.0205third-party2026-08-10
DeepSeek-V4-Pro90.9%±0.0205third-party2026-08-10
Kimi-K2.690.8%±0.0172third-party2026-08-10
Claude Sonnet 590.5%±0.0178third-party2026-08-10
Grok 4.388.8%±0.0196third-party2026-08-10
Kimi-K2.7-Code87.9%±0.0233third-party2026-08-10
GPT-5.4 mini86.9%±0.0241third-party2026-08-10
Qwen3.5-397B-A17B86.4%±0.0245third-party2026-08-10
Qwen3.6-27B85.9%±0.0248third-party2026-08-10
Claude Fable 585.9%±0.0248third-party2026-08-10
GLM-5.185.5%±0.0207third-party2026-08-10
Qwen3.6-35B-A3B84.9%±0.0255third-party2026-08-10
Gemini 3.5 Flash Lite83.3%±0.0266third-party2026-08-10
GLM-4.783.3%unknownthird-party2026-08-10
Qwen3-235B-A22B-Instruct-250780.0%±0.0260third-party2026-08-10
GPT-5.4 nano78.5%±0.0244third-party2026-08-10
DeepSeek-R1-052876.3%unknownthird-party2026-08-10
gemma-4-31B-it75.8%±0.0305third-party2026-08-10
gpt-oss-120b75.8%unknownthird-party2026-08-10
Llama-4-Maverick-17B-128E-Instruct67.0%unknownthird-party2026-08-10
phi-456.1%unknownthird-party2026-08-10
Llama-4-Scout-17B-16E-Instruct51.8%unknownthird-party2026-08-10
Qwen2.5-72B-Instruct49.1%unknownthird-party2026-08-10
Llama-3.3-70B-Instruct47.4%unknownthird-party2026-08-10
gpt-oss-20b46.0%unknownthird-party2026-08-10
Llama-3.1-70B-Instruct44.2%unknownthird-party2026-08-10
gemma-3-27b-it43.9%±0.0354third-party2026-08-10
Mistral-Nemo-Instruct-240729.9%±0.0210third-party2026-08-10
gemma-2-9b-it27.5%unknownthird-party2026-08-10
Llama-3.1-8B-Instruct25.9%unknownthird-party2026-08-10

Epoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. · license CC-BY-4.0 · epoch-ai · https://epoch.ai/benchmarks