Switch language한국어
Back to the list

Google's Gemma 4 AI models get a 3x speed boost by predicting future tokens

TL;DR AI

Key summary

2 min read
  1. Google has released experimental Multi-Token Prediction drafters for Gemma 4 to speed up local AI generation.

  2. The smaller E2B and E4B drafter models predict upcoming tokens and share context with the main model, enabling speculative and sparse decoding.

  3. Google says the approach can deliver up to 3x faster output, with tests showing strong gains on hardware like NVIDIA RTX PRO 6000 GPUs.

  4. Because it keeps inference on-device and Gemma’s license is permissive, the update could make local AI more practical and easier to adopt.

Read the original