Qwen3
Qwen3-32B
Alibaba (Qwen team) · open-permissive apache-2.0
| Parameters (B) | 32.762 |
|---|---|
| Architecture | dense |
| Layers / KV heads / head dim | 64 / 8 / 128 |
| Context window | 40960 |
| Modalities | in: text; out: text |
| GGUF | available |
| Commercial use | yes |
Minimum hardware (estimate)
At Q4_K_M and 8k context, total working set ≈ 22.6 GiB. On a machine with 64 GiB system RAM and no discrete GPU: CPU-only, slow. Limiting factor: No GPU VRAM; fits in system RAM but inference will be CPU-bound.
Scores
| Benchmark | Value | Uncertainty | Run by | Source |
|---|---|---|---|---|
| Aider polyglot | 40.0% | unknown | third-party | epoch-ai |
Provenance
- identity, license, architecture, parameter count: https://huggingface.co/Qwen/Qwen3-32B (retrieved 2026-08-10)
HF: Qwen/Qwen3-32B · snapshot 2026-08-10