Llama 3.1
Llama-3.1-8B-Instruct
Meta AI · open-restricted llama3.1
| Parameters (B) | 8.03 |
|---|---|
| Architecture | dense |
| Layers / KV heads / head dim | 32 / 8 / 128 |
| Context window | 131072 |
| Modalities | in: text; out: text |
| GGUF | available |
| Commercial use | yes |
Minimum hardware (estimate)
At Q4_K_M and 8k context, total working set ≈ 6.5 GiB. On a machine with 64 GiB system RAM and no discrete GPU: CPU-only, slow. Limiting factor: No GPU VRAM; fits in system RAM but inference will be CPU-bound.
Scores
| Benchmark | Value | Uncertainty | Run by | Source |
|---|---|---|---|---|
| GPQA Diamond | 25.9% | unknown | third-party | epoch-ai |
| MATH Level 5 | 22.9% | unknown | third-party | epoch-ai |
| OTIS Mock AIME 2024-2025 | 2.5% | unknown | third-party | epoch-ai |
Provenance
- identity, license, architecture, parameter count: https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct (retrieved 2026-08-10)
HF: meta-llama/Llama-3.1-8B-Instruct · snapshot 2026-08-10