NVIDIA RTX Pro (Blackwell)
Figures assume this GPU plus 32 GB of system RAM (a typical desktop pairing). "What runs on it" is judged at a 8,192-token context. Speeds are estimates, not measurements.
| Model | Sweet-spot quant | Est. speed | Community |
|---|---|---|---|
| DeepSeek-R1-Distill-Llama-8B deepseek-ai | Q8_0 Runs fully on GPU @ 8K ctx | 112–149 tok/s (estimate) | no community data |
| DeepSeek-R1-Distill-Qwen-1.5B deepseek-ai | Q8_0 Runs fully on GPU @ 8K ctx | 505–673 tok/s (estimate) | no community data |
| DeepSeek-R1-Distill-Qwen-14B deepseek-ai | Q8_0 Runs fully on GPU @ 8K ctx | 62–83 tok/s (estimate) | no community data |
| DeepSeek-R1-Distill-Qwen-32B deepseek-ai | Q8_0 Runs fully on GPU @ 8K ctx | 29–39 tok/s (estimate) | no community data |
| DeepSeek-R1-Distill-Qwen-7B deepseek-ai | Q8_0 Runs fully on GPU @ 8K ctx | 125–167 tok/s (estimate) | no community data |
| Devstral-Small-2-24B-Instruct-2512 mistralai | Q8_0 Runs fully on GPU @ 8K ctx | 41–54 tok/s (estimate) | no community data |
| GLM-4.7-Flash zai-org | Q8_0 Runs fully on GPU @ 8K ctx | no community data | |
| Hunyuan-A13B-Instruct tencent | Q8_0 Runs fully on GPU @ 8K ctx | no community data | |
| Kimi-Dev-72B moonshotai | UD-Q5_K_XL Runs fully on GPU @ 8K ctx | 19–25 tok/s (estimate) | no community data |
| Kimi-VL-A3B-Instruct moonshotai | BF16 Runs fully on GPU @ 8K ctx | no community data | |
| Llama-3.1-8B-Instruct meta-llama | Q8_0 Runs fully on GPU @ 8K ctx | 112–149 tok/s (estimate) | no community data |
| Llama-3.3-70B-Instruct meta-llama | Q8_0 Runs fully on GPU @ 8K ctx | 14–18 tok/s (estimate) | no community data |
| Mistral-Small-3.2-24B-Instruct-2506 mistralai | Q8_0 Runs fully on GPU @ 8K ctx | 41–54 tok/s (estimate) | no community data |
| Phi-4-mini-instruct microsoft | Q8_0 Runs fully on GPU @ 8K ctx | 208–278 tok/s (estimate) | no community data |
| Phi-4-reasoning microsoft | Q8_0 Runs fully on GPU @ 8K ctx | 62–83 tok/s (estimate) | no community data |
| Qwen2.5-7B-Instruct Qwen | Q8_0 Runs fully on GPU @ 8K ctx | 125–167 tok/s (estimate) | no community data |
| Qwen2.5-Omni-7B Qwen | BF16 Runs fully on GPU @ 8K ctx | 47–63 tok/s (estimate) | no community data |
| Qwen2.5-VL-32B-Instruct Qwen | Q8_0 Runs fully on GPU @ 8K ctx | 29–39 tok/s (estimate) | no community data |
| Qwen3-14B Qwen | Q8_0 Runs fully on GPU @ 8K ctx | 63–84 tok/s (estimate) | no community data |
| Qwen3-235B-A22B-Instruct-2507 Qwen | UD-Q2_K_XL Runs fully on GPU @ 8K ctx | no community data | |
| Qwen3-30B-A3B-Instruct-2507 Qwen | Q8_0 Runs fully on GPU @ 8K ctx | no community data | |
| Qwen3-32B Qwen | Q8_0 Runs fully on GPU @ 8K ctx | 29–39 tok/s (estimate) | no community data |
| Qwen3-8B Qwen | Q8_0 Runs fully on GPU @ 8K ctx | 108–145 tok/s (estimate) | no community data |
| Qwen3-Coder-30B-A3B-Instruct Qwen | Q8_0 Runs fully on GPU @ 8K ctx | no community data | |
| Qwen3-Embedding-0.6B Qwen | Q8_0 Runs fully on GPU @ 8K ctx | 681–908 tok/s (estimate) | no community data |
| Qwen3-Embedding-4B Qwen | Q4_K_M Runs fully on GPU @ 8K ctx | 290–387 tok/s (estimate) | no community data |
| Qwen3-Embedding-8B Qwen | Q4_K_M Runs fully on GPU @ 8K ctx | 183–244 tok/s (estimate) | no community data |
| Qwen3-Omni-30B-A3B-Instruct Qwen | Q4_K_M Runs fully on GPU @ 8K ctx | no community data | |
| Qwen3-Reranker-0.6B Qwen | BF16 Runs fully on GPU @ 8K ctx | 505–673 tok/s (estimate) | no community data |
| Qwen3-Reranker-4B Qwen | BF16 Runs fully on GPU @ 8K ctx | 116–155 tok/s (estimate) | no community data |
| Qwen3-Reranker-8B Qwen | BF16 Runs fully on GPU @ 8K ctx | 61–82 tok/s (estimate) | no community data |
| Qwen3.6-27B Qwen | Q8_0 Runs fully on GPU @ 8K ctx | 35–47 tok/s (estimate) | no community data |
| Qwen3.6-35B-A3B Qwen | Q8_0 Runs fully on GPU @ 8K ctx | no community data | |
| SmolLM3-3B HuggingFaceTB | BF16 Runs fully on GPU @ 8K ctx | 159–212 tok/s (estimate) | no community data |
| dots.ocr rednote-hilab | BF16 Runs fully on GPU @ 8K ctx | 170–227 tok/s (estimate) | no community data |
| gemma-4-12B-it google | Q8_0 Runs fully on GPU @ 8K ctx | 79–106 tok/s (estimate) | no community data |
| gemma-4-26B-A4B-it google | Q8_0 Runs fully on GPU @ 8K ctx | no community data | |
| gemma-4-31B-it google | Q8_0 Runs fully on GPU @ 8K ctx | 31–41 tok/s (estimate) | no community data |
| gemma-4-E2B-it google | Q8_0 Runs fully on GPU @ 8K ctx | 210–280 tok/s (estimate) | no community data |
| gemma-4-E4B-it google | Q8_0 Runs fully on GPU @ 8K ctx | 129–172 tok/s (estimate) | no community data |
| gpt-oss-120b openai | F16 Runs fully on GPU @ 8K ctx | no community data | |
| gpt-oss-20b openai | F16 Runs fully on GPU @ 8K ctx | no community data | |
| phi-4 microsoft | Q8_0 Runs fully on GPU @ 8K ctx | 62–83 tok/s (estimate) | no community data |
"Est. speed" is a modelled range labelled estimate (D8) for generation (decode) throughput. "Community" shows the median of approved user-submitted reports on this GPU class only where enough exist — never an estimate. "pp" is measured prompt-processing (ingestion) throughput from approved community reports; rows without a measurement show none.
On this GPU, 12 catalog models run fully on the GPU at an 8,192-token context. The most capable is Qwen3-235B-A22B-Instruct-2507 at UD-Q2_K_XL (needs ~93.1 GiB). Pick a smaller model or a lower quant for more headroom.
Biggest model: Qwen3-235B-A22B-Instruct-2507 at UD-Q2_K_XL
llama-server -m Qwen3-235B-A22B-Instruct-2507-UD-Q2_K_XL-00001-of-00002.gguf -c 8192 -ngl 999
Derived from the fit engine at an 8,192-token context. See more answer packs.
A second RTX Pro 6000 Blackwell 96GB pools VRAM: 2 × 96 GB = 192 GB combined. Bigger models can then load because their weights split across both cards — but a second card does not make generation proportionally faster (see the reality check below).
A second RTX Pro 6000 Blackwell 96GB adds 96 GB of VRAM for about $11,830 — ≈$123.23/GB of added VRAM (new price as of 2026-07-18 — Newegg).
3 more catalog models could newly fit fully in the combined 192 GB at a 8,192-token context — for example:
Estimate — assumes the model's weights split across both cards (a layer split, as llama.cpp does by default). This is a combined-VRAM projection, not a measured or verdict-chipped result: the fit engine treats two cards as one summed memory pool and does not model the link between them. Check your exact model and context in the calculator.
The honest reality of a second card
Read: Multi-GPU for local LLMs — when a second card is worth it →
Derived from the fit engine at a 8,192-token context, comparing a single 96 GB card against a summed 192 GB two-card pool.
Prices are point-in-time observations, not live quotes.
Not enough history yet — trends appear once at least three dated observations are recorded.
selected to compare · pick at least 2