Largest writing-capable models that fit 16 GB

On a 16 GB GPU, 12 catalog models run fully on the GPU at an 8,192-token context. The most capable is gemma-4-26B-A4B-it at QAT-Q4_0 (needs ~15.8 GiB). Pick a smaller model or a lower quant for more headroom.

Figures assume a 16 GB GPU paired with 32 GB of system RAM (a typical desktop), judged at an 8,192-token context. Speed depends on the specific card — this page ranks by fit, not speed.

The biggest model you can run

gemma-4-26B-A4B-it at QAT-Q4_0 · 26.5B params · needs ~15.8 GiB

One command to run it (llama.cpp):

llama-server -m gemma-4-26B_q4_0-it.gguf -c 8192 -ngl 999

Models that run, largest first

Model Sweet-spot quant Fits in
gemma-4-26B-A4B-it google QAT-Q4_0 Runs fully on GPU ~15.8 GiB
Devstral-Small-2-24B-Instruct-2512 mistralai IQ4_XS Runs fully on GPU ~14.9 GiB
Mistral-Small-3.2-24B-Instruct-2506 mistralai IQ4_XS Runs fully on GPU ~14.9 GiB
gpt-oss-20b openai F16 Runs fully on GPU ~15 GiB
Kimi-VL-A3B-Instruct moonshotai Q4_K_M Runs fully on GPU ~11.6 GiB
DeepSeek-R1-Distill-Qwen-14B deepseek-ai Q4_K_M Runs fully on GPU ~11.2 GiB
Qwen3-14B Qwen Q4_K_M Runs fully on GPU ~11 GiB
Phi-4-reasoning microsoft Q4_K_M Runs fully on GPU ~11.4 GiB
phi-4 microsoft Q4_K_M Runs fully on GPU ~11.2 GiB
gemma-4-12B-it google Q8_0 Runs fully on GPU ~14.3 GiB
Qwen3-8B Qwen Q8_0 Runs fully on GPU ~10.6 GiB
DeepSeek-R1-Distill-Llama-8B deepseek-ai Q8_0 Runs fully on GPU ~10.3 GiB

Derived live from the fit engine + catalog at an 8,192-token context. "Fits in" is the modelled VRAM the sweet-spot quant needs (weights + KV cache + overhead). Speed depends on your specific card — check a GPU page or the calculator.

Check it against your exact setup

Open the fit calculator

selected to compare · pick at least 2