AMD Ryzen AI Max+ 395 (128 GB)

Mini AI PC

Unified memory
128 GB
Memory bandwidth
256 GB/s
Class
Minipc

Figures assume the 128 GB unified memory pool shared between CPU and GPU. "What runs on it" is judged at a 8,192-token context. Speeds are estimates, not measurements.

What runs on it

Model Sweet-spot quant Est. speed Community
DeepSeek-R1-Distill-Llama-8B deepseek-ai Q8_0 Runs fully on GPU @ 8K ctx 16–21 tok/s (estimate) no community data
DeepSeek-R1-Distill-Qwen-1.5B deepseek-ai Q8_0 Runs fully on GPU @ 8K ctx 72–96 tok/s (estimate) no community data
DeepSeek-R1-Distill-Qwen-14B deepseek-ai Q8_0 Runs fully on GPU @ 8K ctx 9–12 tok/s (estimate) no community data
DeepSeek-R1-Distill-Qwen-32B deepseek-ai Q8_0 Runs fully on GPU @ 8K ctx 4–6 tok/s (estimate) no community data
DeepSeek-R1-Distill-Qwen-7B deepseek-ai Q8_0 Runs fully on GPU @ 8K ctx 18–24 tok/s (estimate) no community data
Devstral-Small-2-24B-Instruct-2512 mistralai Q8_0 Runs fully on GPU @ 8K ctx 6–8 tok/s (estimate) no community data
GLM-4.7-Flash zai-org Q8_0 Runs fully on GPU @ 8K ctx no community data
Hunyuan-A13B-Instruct tencent Q8_0 Runs fully on GPU @ 8K ctx no community data
Kimi-Dev-72B moonshotai UD-Q5_K_XL Runs fully on GPU @ 8K ctx 3–4 tok/s (estimate) no community data
Kimi-VL-A3B-Instruct moonshotai BF16 Runs fully on GPU @ 8K ctx no community data
Llama-3.1-8B-Instruct meta-llama Q8_0 Runs fully on GPU @ 8K ctx 16–21 tok/s (estimate) no community data
Llama-3.3-70B-Instruct meta-llama Q8_0 Runs fully on GPU @ 8K ctx 2–3 tok/s (estimate) no community data
Mistral-Small-3.2-24B-Instruct-2506 mistralai Q8_0 Runs fully on GPU @ 8K ctx 6–8 tok/s (estimate) no community data
Phi-4-mini-instruct microsoft Q8_0 Runs fully on GPU @ 8K ctx 30–40 tok/s (estimate) no community data
Phi-4-reasoning microsoft Q8_0 Runs fully on GPU @ 8K ctx 9–12 tok/s (estimate) no community data
Qwen2.5-7B-Instruct Qwen Q8_0 Runs fully on GPU @ 8K ctx 18–24 tok/s (estimate) no community data
Qwen2.5-Omni-7B Qwen BF16 Runs fully on GPU @ 8K ctx 7–9 tok/s (estimate) no community data
Qwen2.5-VL-32B-Instruct Qwen Q8_0 Runs fully on GPU @ 8K ctx 4–6 tok/s (estimate) no community data
Qwen3-14B Qwen Q8_0 Runs fully on GPU @ 8K ctx 9–12 tok/s (estimate) no community data
Qwen3-235B-A22B-Instruct-2507 Qwen UD-Q2_K_XL Runs fully on GPU @ 8K ctx no community data
Qwen3-30B-A3B-Instruct-2507 Qwen Q8_0 Runs fully on GPU @ 8K ctx no community data
Qwen3-32B Qwen Q8_0 Runs fully on GPU @ 8K ctx 4–6 tok/s (estimate) no community data
Qwen3-8B Qwen Q8_0 Runs fully on GPU @ 8K ctx 15–21 tok/s (estimate) no community data
Qwen3-Coder-30B-A3B-Instruct Qwen Q8_0 Runs fully on GPU @ 8K ctx no community data
Qwen3-Embedding-0.6B Qwen Q8_0 Runs fully on GPU @ 8K ctx 97–130 tok/s (estimate) no community data
Qwen3-Embedding-4B Qwen Q4_K_M Runs fully on GPU @ 8K ctx 41–55 tok/s (estimate) no community data
Qwen3-Embedding-8B Qwen Q4_K_M Runs fully on GPU @ 8K ctx 26–35 tok/s (estimate) no community data
Qwen3-Omni-30B-A3B-Instruct Qwen Q4_K_M Runs fully on GPU @ 8K ctx no community data
Qwen3-Reranker-0.6B Qwen BF16 Runs fully on GPU @ 8K ctx 72–96 tok/s (estimate) no community data
Qwen3-Reranker-4B Qwen BF16 Runs fully on GPU @ 8K ctx 17–22 tok/s (estimate) no community data
Qwen3-Reranker-8B Qwen BF16 Runs fully on GPU @ 8K ctx 9–12 tok/s (estimate) no community data
Qwen3.6-27B Qwen Q8_0 Runs fully on GPU @ 8K ctx 5–7 tok/s (estimate) no community data
Qwen3.6-35B-A3B Qwen Q8_0 Runs fully on GPU @ 8K ctx no community data
SmolLM3-3B HuggingFaceTB BF16 Runs fully on GPU @ 8K ctx 23–30 tok/s (estimate) no community data
dots.ocr rednote-hilab BF16 Runs fully on GPU @ 8K ctx 24–32 tok/s (estimate) no community data
gemma-4-12B-it google Q8_0 Runs fully on GPU @ 8K ctx 11–15 tok/s (estimate) no community data
gemma-4-26B-A4B-it google Q8_0 Runs fully on GPU @ 8K ctx no community data
gemma-4-31B-it google Q8_0 Runs fully on GPU @ 8K ctx 4–6 tok/s (estimate) no community data
gemma-4-E2B-it google Q8_0 Runs fully on GPU @ 8K ctx 30–40 tok/s (estimate) no community data
gemma-4-E4B-it google Q8_0 Runs fully on GPU @ 8K ctx 18–25 tok/s (estimate) no community data
gpt-oss-120b openai F16 Runs fully on GPU @ 8K ctx no community data
gpt-oss-20b openai F16 Runs fully on GPU @ 8K ctx no community data
phi-4 microsoft Q8_0 Runs fully on GPU @ 8K ctx 9–12 tok/s (estimate) no community data

"Est. speed" is a modelled range labelled estimate (D8) for generation (decode) throughput. "Community" shows the median of approved user-submitted reports on this GPU class only where enough exist — never an estimate. "pp" is measured prompt-processing (ingestion) throughput from approved community reports; rows without a measurement show none.

What can I run on a AMD Ryzen AI Max+ 395 (128 GB)?

On this GPU, 12 catalog models run fully on the GPU at an 8,192-token context. The most capable is Qwen3-235B-A22B-Instruct-2507 at UD-Q2_K_XL (needs ~93.1 GiB). Pick a smaller model or a lower quant for more headroom.

Biggest model: Qwen3-235B-A22B-Instruct-2507 at UD-Q2_K_XL

llama-server -m Qwen3-235B-A22B-Instruct-2507-UD-Q2_K_XL-00001-of-00002.gguf -c 8192 -ngl 999

Derived from the fit engine at an 8,192-token context. See more answer packs.

selected to compare · pick at least 2