Llama 3.1

Llama-3.1-8B-Instruct

Meta AI · open-restricted llama3.1

Parameters (B)8.03
Architecturedense
Layers / KV heads / head dim32 / 8 / 128
Context window131072
Modalitiesin: text; out: text
GGUFavailable
Commercial useyes

Minimum hardware (estimate)

At Q4_K_M and 8k context, total working set ≈ 6.5 GiB. On a machine with 64 GiB system RAM and no discrete GPU: CPU-only, slow. Limiting factor: No GPU VRAM; fits in system RAM but inference will be CPU-bound.

Open in fit advisor

Scores

BenchmarkValueUncertaintyRun bySource
GPQA Diamond25.9%unknownthird-partyepoch-ai
MATH Level 522.9%unknownthird-partyepoch-ai
OTIS Mock AIME 2024-20252.5%unknownthird-partyepoch-ai

Provenance

HF: meta-llama/Llama-3.1-8B-Instruct · snapshot 2026-08-10