Google AI Releases Multi-Token Prediction (MTP) Drafters for Gemma 4: Delivering Up to 3x Faster Inference Without Quality Loss

TL;DR AI
2 min readKey summary
Google released Multi-Token Prediction (MTP) drafters for Gemma 4 to speed up inference.
The approach uses speculative decoding: a small drafter proposes several tokens, and the larger model verifies them in one pass.
Google says this can cut generation latency by up to 3x while preserving exact output quality.
The update also includes extra optimizations for edge variants, helping reduce serving cost and improve responsiveness.
