Gemma 4

gemma-4-31B-it

Google DeepMind · open-permissive apache-2.0

Parameters (B)31.273
Architecturedense
Layers / KV heads / head dim60 / 16 / 256
Context window262144
Modalitiesin: text, image; out: text
GGUFavailable
Commercial useyes

Minimum hardware (estimate)

At Q4_K_M and 8k context, total working set ≈ 27.6 GiB. On a machine with 64 GiB system RAM and no discrete GPU: CPU-only, slow. Limiting factor: No GPU VRAM; fits in system RAM but inference will be CPU-bound.

Open in fit advisor

Scores

BenchmarkValueUncertaintyRun bySource
GPQA Diamond75.8%±0.0305third-partyepoch-ai
OTIS Mock AIME 2024-202573.3%±0.0667third-partyepoch-ai
SimpleQA Verified9.6%±0.0094third-partyepoch-ai

Provenance

HF: google/gemma-4-31B-it · snapshot 2026-08-10