Gemma 4

gemma-4-12B-it

Google DeepMind · open-permissive apache-2.0

Parameters (B)11.96
Architecturedense
Layers / KV heads / head dim48 / 8 / 256
Context window262144
Modalitiesin: text, image, audio; out: text
GGUFavailable
Commercial useyes

Minimum hardware (estimate)

At Q4_K_M and 8k context, total working set ≈ 11.0 GiB. On a machine with 64 GiB system RAM and no discrete GPU: CPU-only, slow. Limiting factor: No GPU VRAM; fits in system RAM but inference will be CPU-bound.

Open in fit advisor

Scores

No Tier A comparable scores ingested yet for this model.

Provenance

HF: google/gemma-4-12B-it · snapshot 2026-08-10