SEGA: Spectral-Energy Guided Attention for Resolution Extrapolation in Diffusion Transformers
TL;DR AI
2 min readKey summary
SEGA is a training-free method for high-resolution text-to-image generation with diffusion transformers.
It uses the latent’s spatial-frequency structure to adaptively scale attention across RoPE components during denoising.
This helps diffusion transformers extrapolate beyond their training resolution range while preserving global structure and fine detail.
The result is better high-resolution synthesis than prior training-free baselines, without extra training.
