Google made Gemma 4 models 3x faster with MTP Drafters

TL;DR AI
2 min readKey summary
Google added speculative decoding to Gemma 4, speeding up inference on consumer GPUs and edge devices.
The new setup uses a smaller MTP drafter to suggest tokens while the larger model verifies them in parallel.
Gemma 4 26B MoE and 31B Dense can run up to 3x faster without changing output quality.
The update lowers latency for local AI apps like chatbots, coding assistants, and autonomous agents.
