SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
TL;DR AI
2 min readKey summary
Researchers unveiled SANA-Video 2.0, a hybrid video diffusion transformer for efficient video generation.
The 5B and 14B models combine linear and softmax attention, plus attention residuals, to improve token interaction.
Trained from scratch, the system can generate up to 720p video efficiently on a single GPU.
The approach aims to keep much of softmax attention’s quality while preserving the scalability of linear attention.
