Which quant should I pick?

Pick your GPU and a model, and we'll rank its quants from highest to lowest quality — with the fit verdict, size, an estimated speed, and a plain-English note on what you give up at each step down. No account needed.

1 · Your GPU

  • No matching GPU in the list — you can type the model manually.

Search for your card above (Apple Silicon routes to unified memory automatically). Without a GPU we'll still show sizes and the quality guidance below.

Used to judge partial CPU offload when a quant overflows VRAM.

Save your rig for one-click answers.

Your rig lives in this browser. A free account keeps it across the calculator, model pages and this helper.

Save my rig

2 · Model & context

Lower precision shrinks the KV cache (fits more context).

Phi-4-mini-instruct quants, ranked

Full model page & downloads →
Pick your GPU above to see which of these fit, plus estimated speeds. Sizes and the quality guidance below don't need a rig.
  1. GGUF · Q8_0
    3.80 GB

    Effectively lossless in community experience — the closest you get to the original model, at the largest practical GGUF size.

    Note: Bartowski Q8_0 quant of Phi-4-mini-instruct (3.836B dense) — high-quality option. MIT.

  2. GGUF · Q4_K_M Community default
    2.32 GB

    The community default: minor quality loss versus Q5/Q6, usually unnoticeable outside edge tasks, at a much smaller size.

    Note: Balanced size/quality — good default.

Verdicts and sizes match this model's page exactly. Speeds are roofline estimates (a range, not a promise). Quality notes reflect community experience, not measured benchmarks.