GLM-4
GLM-4.7-Flash
Z.ai (Zhipu AI) · open-permissive mit
| Parameters (B) | 31.221 |
|---|---|
| Architecture | moe |
| Layers / KV heads / head dim | 47 / 20 / 102 |
| Context window | 202752 |
| Modalities | in: text; out: text |
| GGUF | available |
| Commercial use | yes |
Minimum hardware (estimate)
At Q4_K_M and 8k context, total working set ≈ 22.6 GiB. On a machine with 64 GiB system RAM and no discrete GPU: CPU-only, slow. Limiting factor: No GPU VRAM; fits in system RAM but inference will be CPU-bound.
Scores
No Tier A comparable scores ingested yet for this model.
Provenance
- identity, license, architecture, parameter count: https://huggingface.co/zai-org/GLM-4.7-Flash (retrieved 2026-08-10)
HF: zai-org/GLM-4.7-Flash · snapshot 2026-08-10