Which quant should I pick?

Pick your GPU and a model, and we'll rank its quants from highest to lowest quality — with the fit verdict, size, an estimated speed, and a plain-English note on what you give up at each step down. No account needed.

1 · Your GPU

  • No matching GPU in the list — you can type the model manually.

Search for your card above (Apple Silicon routes to unified memory automatically). Without a GPU we'll still show sizes and the quality guidance below.

Used to judge partial CPU offload when a quant overflows VRAM.

Save your rig for one-click answers.

Your rig lives in this browser. A free account keeps it across the calculator, model pages and this helper.

Save my rig

2 · Model & context

Lower precision shrinks the KV cache (fits more context).

GLM-4.7 quants, ranked

Full model page & downloads →
Pick your GPU above to see which of these fit, plus estimated speeds. Sizes and the quality guidance below don't need a rig.
  1. GGUF · Q4_K_M Community default
    201.59 GB

    The community default: minor quality loss versus Q5/Q6, usually unnoticeable outside edge tasks, at a much smaller size.

    Note: Unsloth Q4_K_M quant of the GLM-4.7 358B-A32B MoE giant — preservation copy. MIT (Z.ai).

  2. GGUF · IQ4_XS
    178.41 GB

    An i-quant that packs 4-bit weights tighter than Q4_K_S — similar quality in community reports and a smaller file, though slightly heavier to decode on some runtimes.

    Note: Unsloth IQ4_XS (imatrix 4-bit) quant of the GLM-4.7 358B-A32B MoE giant — preservation copy. MIT (Z.ai).

  3. GGUF · UD-Q2_K_XL
    125.91 GB

    Note: Unsloth dynamic 2-bit (UD-Q2_K_XL) — accessible 358B-A32B MoE giant lead (~126 GB, 3 shards; preservation copy). MIT.

Verdicts and sizes match this model's page exactly. Speeds are roofline estimates (a range, not a promise). Quality notes reflect community experience, not measured benchmarks.