Switch language한국어
Back to the list

Moonshot AI Open-Sources FlashKDA: CUTLASS Kernels for Kimi Delta Attention with Variable-Length Batching and H20 Benchmarks

TL;DR AI

Key summary

2 min read
  1. Moonshot AI open-sourced FlashKDA under an MIT license as a CUTLASS-based CUDA kernel for Kimi Delta Attention.

  2. The company says it delivers 1.72× to 2.22× faster prefill than flash-linear-attention on NVIDIA H20 GPUs.

  3. FlashKDA also adds variable-length batching support, making deployment easier for real-world workloads.

  4. The release strengthens Moonshot AI’s Kimi Linear stack for more efficient long-context LLM inference.

Read the original