Switch language한국어
Back to the list

Google announces "multi-token prediction," a technology that uses a small AI to generate drafts and speed up large AI

TL;DR AI

Key summary

2 min read
  1. Google introduced Multi-token Prediction for Gemma 4, a speculative decoding approach that lets a small drafter propose several tokens ahead.

  2. The main model verifies those tokens in parallel, reducing inference latency while preserving output quality.

  3. Google says the method delivered up to about 3x speedups in tests on some hardware, including NVIDIA RTX PRO 6000, Pixel TPU, and NVIDIA A100.

  4. Draft models for several Gemma 4 variants are now available on Hugging Face.

Read the original