Apple M3
Figures assume the 96 GB unified memory pool shared between CPU and GPU. "What runs on it" is judged at a 8,192-token context. Speeds are estimates, not measurements.
| Model | Sweet-spot quant | Est. speed | Community |
|---|---|---|---|
| DeepSeek-R1-Distill-Llama-8B deepseek-ai | Q8_0 Runs fully on GPU @ 8K ctx | 51–68 tok/s (estimate) | no community data |
| DeepSeek-R1-Distill-Qwen-1.5B deepseek-ai | Q8_0 Runs fully on GPU @ 8K ctx | 231–308 tok/s (estimate) | no community data |
| DeepSeek-R1-Distill-Qwen-14B deepseek-ai | Q8_0 Runs fully on GPU @ 8K ctx | 28–38 tok/s (estimate) | no community data |
| DeepSeek-R1-Distill-Qwen-32B deepseek-ai | Q8_0 Runs fully on GPU @ 8K ctx | 13–18 tok/s (estimate) | no community data |
| DeepSeek-R1-Distill-Qwen-7B deepseek-ai | Q8_0 Runs fully on GPU @ 8K ctx | 57–76 tok/s (estimate) | no community data |
| Devstral-Small-2-24B-Instruct-2512 mistralai | Q8_0 Runs fully on GPU @ 8K ctx | 19–25 tok/s (estimate) | no community data |
| GLM-4.7-Flash zai-org | Q8_0 Runs fully on GPU @ 8K ctx | no community data | |
| Hunyuan-A13B-Instruct tencent | Q5_K_M Runs fully on GPU @ 8K ctx | no community data | |
| Kimi-Dev-72B moonshotai | UD-Q5_K_XL Runs fully on GPU @ 8K ctx | 9–12 tok/s (estimate) | no community data |
| Kimi-VL-A3B-Instruct moonshotai | BF16 Runs fully on GPU @ 8K ctx | no community data | |
| Llama-3.1-8B-Instruct meta-llama | Q8_0 Runs fully on GPU @ 8K ctx | 51–68 tok/s (estimate) | no community data |
| Llama-3.3-70B-Instruct meta-llama | Q4_K_M Runs fully on GPU @ 8K ctx | 11–14 tok/s (estimate) | no community data |
| Mistral-Small-3.2-24B-Instruct-2506 mistralai | Q8_0 Runs fully on GPU @ 8K ctx | 19–25 tok/s (estimate) | no community data |
| Phi-4-mini-instruct microsoft | Q8_0 Runs fully on GPU @ 8K ctx | 95–127 tok/s (estimate) | no community data |
| Phi-4-reasoning microsoft | Q8_0 Runs fully on GPU @ 8K ctx | 28–38 tok/s (estimate) | no community data |
| Qwen2.5-7B-Instruct Qwen | Q8_0 Runs fully on GPU @ 8K ctx | 57–76 tok/s (estimate) | no community data |
| Qwen2.5-Omni-7B Qwen | BF16 Runs fully on GPU @ 8K ctx | 22–29 tok/s (estimate) | no community data |
| Qwen2.5-VL-32B-Instruct Qwen | Q8_0 Runs fully on GPU @ 8K ctx | 13–18 tok/s (estimate) | no community data |
| Qwen3-14B Qwen | Q8_0 Runs fully on GPU @ 8K ctx | 29–38 tok/s (estimate) | no community data |
| Qwen3-30B-A3B-Instruct-2507 Qwen | Q8_0 Runs fully on GPU @ 8K ctx | no community data | |
| Qwen3-32B Qwen | Q8_0 Runs fully on GPU @ 8K ctx | 13–18 tok/s (estimate) | no community data |
| Qwen3-8B Qwen | Q8_0 Runs fully on GPU @ 8K ctx | 50–66 tok/s (estimate) | no community data |
| Qwen3-Coder-30B-A3B-Instruct Qwen | Q8_0 Runs fully on GPU @ 8K ctx | no community data | |
| Qwen3-Embedding-0.6B Qwen | Q8_0 Runs fully on GPU @ 8K ctx | 311–415 tok/s (estimate) | no community data |
| Qwen3-Embedding-4B Qwen | Q4_K_M Runs fully on GPU @ 8K ctx | 133–177 tok/s (estimate) | no community data |
| Qwen3-Embedding-8B Qwen | Q4_K_M Runs fully on GPU @ 8K ctx | 84–111 tok/s (estimate) | no community data |
| Qwen3-Omni-30B-A3B-Instruct Qwen | Q4_K_M Runs fully on GPU @ 8K ctx | no community data | |
| Qwen3-Reranker-0.6B Qwen | BF16 Runs fully on GPU @ 8K ctx | 231–307 tok/s (estimate) | no community data |
| Qwen3-Reranker-4B Qwen | BF16 Runs fully on GPU @ 8K ctx | 53–71 tok/s (estimate) | no community data |
| Qwen3-Reranker-8B Qwen | BF16 Runs fully on GPU @ 8K ctx | 28–37 tok/s (estimate) | no community data |
| Qwen3.6-27B Qwen | Q8_0 Runs fully on GPU @ 8K ctx | 16–21 tok/s (estimate) | no community data |
| Qwen3.6-35B-A3B Qwen | Q8_0 Runs fully on GPU @ 8K ctx | no community data | |
| SmolLM3-3B HuggingFaceTB | BF16 Runs fully on GPU @ 8K ctx | 73–97 tok/s (estimate) | no community data |
| dots.ocr rednote-hilab | BF16 Runs fully on GPU @ 8K ctx | 78–104 tok/s (estimate) | no community data |
| gemma-4-12B-it google | Q8_0 Runs fully on GPU @ 8K ctx | 36–48 tok/s (estimate) | no community data |
| gemma-4-26B-A4B-it google | Q8_0 Runs fully on GPU @ 8K ctx | no community data | |
| gemma-4-31B-it google | Q8_0 Runs fully on GPU @ 8K ctx | 14–19 tok/s (estimate) | no community data |
| gemma-4-E2B-it google | Q8_0 Runs fully on GPU @ 8K ctx | 96–128 tok/s (estimate) | no community data |
| gemma-4-E4B-it google | Q8_0 Runs fully on GPU @ 8K ctx | 59–78 tok/s (estimate) | no community data |
| gpt-oss-120b openai | F16 Runs fully on GPU @ 8K ctx | no community data | |
| gpt-oss-20b openai | F16 Runs fully on GPU @ 8K ctx | no community data | |
| phi-4 microsoft | Q8_0 Runs fully on GPU @ 8K ctx | 28–38 tok/s (estimate) | no community data |
"Est. speed" is a modelled range labelled estimate (D8) for generation (decode) throughput. "Community" shows the median of approved user-submitted reports on this GPU class only where enough exist — never an estimate. "pp" is measured prompt-processing (ingestion) throughput from approved community reports; rows without a measurement show none.
On this GPU, 12 catalog models run fully on the GPU at an 8,192-token context. The most capable is gpt-oss-120b at F16 (needs ~68.2 GiB). Pick a smaller model or a lower quant for more headroom.
Biggest model: gpt-oss-120b at F16
llama-server -m gpt-oss-120b-F16.gguf -c 8192 -ngl 999
Derived from the fit engine at an 8,192-token context. See more answer packs.
selected to compare · pick at least 2