PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
TL;DR AI
2 min readKey summary
PiD introduces a pixel diffusion decoder for latent-to-image synthesis in pixel space.
It uses sigma-aware latent conditioning and DMD2 distillation to cut inference to just four steps.
The approach upscales VAE and semantic latents into megapixel images with lower latency and memory use.
It aims to improve visual fidelity over cascaded super-resolution pipelines on high-end GPUs like RTX 5090 and GB200.
