GLM-4

GLM-4.7-Flash

Z.ai (Zhipu AI) · open-permissive mit

Parameters (B)31.221
Architecturemoe
Layers / KV heads / head dim47 / 20 / 102
Context window202752
Modalitiesin: text; out: text
GGUFavailable
Commercial useyes

Minimum hardware (estimate)

At Q4_K_M and 8k context, total working set ≈ 22.6 GiB. On a machine with 64 GiB system RAM and no discrete GPU: CPU-only, slow. Limiting factor: No GPU VRAM; fits in system RAM but inference will be CPU-bound.

Open in fit advisor

Scores

No Tier A comparable scores ingested yet for this model.

Provenance

HF: zai-org/GLM-4.7-Flash · snapshot 2026-08-10