Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding
TL;DR AI
2 min readKey summary
Researchers introduced Graft, a lossless training-free framework for speculative decoding in large language models.
Graft dynamically combines pruning with token retrieval to recover draft-tree coverage and improve acceptance rates.
The method speeds up inference while preserving output quality, making speculative decoding more practical.
It showed gains across multiple benchmarks and contexts, including comparisons with EAGLE-3, Qwen3-235B, and DFlash.
