Accelerating Gemma 4: faster inference with multi-token prediction drafters
TL;DR AI
2 min readKey summary
Google’s Gemma 4 update puts multi-token prediction and speculative decoding front and center as a way to speed up inference.
Hacker News commenters said the approach can deliver big throughput gains, but real-world behavior depends on the serving stack and model setup.
Some users reported compatibility and tooling quirks in projects like LM Studio and llama-server, especially around quantization and function calling.
The discussion also compared Gemma’s speed and quality with Qwen, Gemini, and other open or commercial models, highlighting practical tradeoffs for deployment.



