Google speeds up Gemma 4 threefold with multi-token prediction

TL;DR AI
2 min readKey summary
Google added multi-token prediction drafters to Gemma 4 to speed up text generation by up to 3x.
A small auxiliary model proposes several tokens ahead, and the main model verifies them in one pass.
The technique reduces LLM inference bottlenecks without changing output quality.
It could make open-weight AI models more practical on phones, PCs, and cloud systems.
