Switch language한국어
Back to the list

Google made Gemma 4 models 3x faster with MTP Drafters

TL;DR AI

Key summary

2 min read
  1. Google added speculative decoding to Gemma 4, speeding up inference on consumer GPUs and edge devices.

  2. The new setup uses a smaller MTP drafter to suggest tokens while the larger model verifies them in parallel.

  3. Gemma 4 26B MoE and 31B Dense can run up to 3x faster without changing output quality.

  4. The update lowers latency for local AI apps like chatbots, coding assistants, and autonomous agents.

Read the original