Which quant should I pick?

Pick your GPU and a model, and we'll rank its quants from highest to lowest quality — with the fit verdict, size, an estimated speed, and a plain-English note on what you give up at each step down. No account needed.

1 · Your GPU

  • No matching GPU in the list — you can type the model manually.

Search for your card above (Apple Silicon routes to unified memory automatically). Without a GPU we'll still show sizes and the quality guidance below.

Used to judge partial CPU offload when a quant overflows VRAM.

Save your rig for one-click answers.

Your rig lives in this browser. A free account keeps it across the calculator, model pages and this helper.

Save my rig

2 · Model & context

Lower precision shrinks the KV cache (fits more context).

Mistral-Small-3.2-24B-Instruct-2506 quants, ranked

Full model page & downloads →
Pick your GPU above to see which of these fit, plus estimated speeds. Sizes and the quality guidance below don't need a rig.
  1. GGUF · Q8_0
    23.33 GB

    Effectively lossless in community experience — the closest you get to the original model, at the largest practical GGUF size.

    Note: Bartowski Q8_0 quant of Mistral-Small-3.2-24B-Instruct-2506 (24.011B dense) — high-quality option. Apache-2.0.

  2. GGUF · Q4_K_M Community default
    13.35 GB

    The community default: minor quality loss versus Q5/Q6, usually unnoticeable outside edge tasks, at a much smaller size.

    Note: Balanced size/quality — good default.

  3. GGUF · IQ4_XS
    11.88 GB

    An i-quant that packs 4-bit weights tighter than Q4_K_S — similar quality in community reports and a smaller file, though slightly heavier to decode on some runtimes.

    Note: Bartowski IQ4_XS quant of Mistral-Small-3.2-24B-Instruct-2506 (24.011B dense) — compact low-bit option. Apache-2.0.

Verdicts and sizes match this model's page exactly. Speeds are roofline estimates (a range, not a promise). Quality notes reflect community experience, not measured benchmarks.