Gemma 4 on a DGX Spark with vLLM: MTP speculative decoding on NVFP4 took decode from 33.8 to 14.5 ms/token and 8-stream output from 179 to 306 tok/s. Full write-up: www.agileguy.ca/tuning-gemma-4-on-a…
Technologist seeking to improve the human experience daemon.agileguy.ca