Qwen
Community-submitted & scraped throughput measurements. Sortable by generation speed.
← Back to model| Rig / GPU | Quant | Runtime | Context | Prompt tok/s | Gen tok/s | VRAM (GB) | Source | Reporter |
|---|---|---|---|---|---|---|---|---|
| NVIDIA DGX Spark (128 GB) | GGUF · Q8_0 | llama.cpp | 2,048 | 1,654.25 | 44.26 | — | scraped | anonymous |
| NVIDIA DGX Spark (128 GB) | GGUF · Q8_0 | llama.cpp | 32,768 | 686.45 | 26.92 | — | scraped | anonymous |
selected to compare · pick at least 2