Gemma 2
gemma-2-9b-it
Google DeepMind · open-restricted gemma
| Parameters (B) | 9.242 |
|---|---|
| Architecture | dense |
| Layers / KV heads / head dim | 42 / 8 / 256 |
| Context window | 8192 |
| Modalities | in: text; out: text |
| GGUF | available |
| Commercial use | yes |
Minimum hardware (estimate)
At Q4_K_M and 8k context, total working set ≈ 9.0 GiB. On a machine with 64 GiB system RAM and no discrete GPU: CPU-only, slow. Limiting factor: No GPU VRAM; fits in system RAM but inference will be CPU-bound.
Scores
| Benchmark | Value | Uncertainty | Run by | Source |
|---|---|---|---|---|
| GPQA Diamond | 27.5% | unknown | third-party | epoch-ai |
| MATH Level 5 | 21.0% | unknown | third-party | epoch-ai |
| OTIS Mock AIME 2024-2025 | 0.6% | unknown | third-party | epoch-ai |
Provenance
- identity, license, architecture, parameter count: https://huggingface.co/google/gemma-2-9b-it (retrieved 2026-08-10)
HF: google/gemma-2-9b-it · snapshot 2026-08-10