DFSAttn: Dynamic Fine-grained Sparse Attention for Efficient Video Generation

TL;DR AI
2 min readKey summary
Researchers introduced DFSAttn, a training-free sparse attention method for diffusion transformers in video generation.
DFSAttn combines token reordering, hierarchical block scoring, and sparse mask caching with adaptive ratios for fine-grained sparsity.
The method cuts the high cost of full spatiotemporal attention and delivers up to 2.1x end-to-end speedup.
It also maintains or improves quality versus prior sparse-attention methods, especially at high sparsity levels.
