google

gemma-4-31B-it — performance reports

Community-submitted & scraped throughput measurements. Sortable by generation speed.

← Back to model
Ran this model? Log in to submit your own performance result.
Rig / GPU Quant Runtime Context Prompt tok/s Gen tok/s VRAM (GB) Source Reporter
RTX 3090 GGUF · Q4_K_M llama.cpp 4,096 1,155.80 34.70 scraped anonymous
RTX 3090 GGUF · Q4_K_M llama.cpp 16,384 913.20 33.50 scraped anonymous
RTX 3090 GGUF · Q4_K_M llama.cpp 32,768 723.70 31.40 scraped anonymous

selected to compare · pick at least 2