Sakana AI and NVIDIA Introduce TwELL with CUDA Kernels for 20.5% Inference and 21.9% Training Speedup in LLMs

TL;DR AI
2 min readKey summary
Sakana AI and NVIDIA introduced TwELL, a tile-based sparse representation for LLM feedforward layers.
Integrated directly into CUDA kernels, TwELL avoids conversion overhead and better exploits activation sparsity on GPUs.
The approach improves batched performance, with reported speedups of 20.5% for inference and 21.9% for training.
It offers a practical path to cheaper high-throughput LLM workloads without changing model architecture.
