Got Gemma 4 running locally on CUDA, first and fastest local inference to my knowledge.
RTX 3090:
• BF16 float: 110 tok/s
• Q4_K_M GGUF: 170 tok/s
Both quantized and full precision, token-for-token match with @hf.co transformers.
CEO @ Studio Atelico | Chief Scientist @ Predibase | Ex-UberAI, Stanford