Gemma 4

gemma-4-E4B-it

Google DeepMind · open-permissive apache-2.0

Parameters (B)7.996
Architecturedense
Layers / KV heads / head dim42 / 2 / 256
Context window131072
Modalitiesin: text, image, audio; out: text
GGUFavailable
Commercial useyes

Minimum hardware (estimate)

At Q4_K_M and 8k context, total working set ≈ 6.1 GiB. On a machine with 64 GiB system RAM and no discrete GPU: CPU-only, slow. Limiting factor: No GPU VRAM; fits in system RAM but inference will be CPU-bound.

Open in fit advisor

Scores

No Tier A comparable scores ingested yet for this model.

Provenance

HF: google/gemma-4-E4B-it · snapshot 2026-08-10