JLT: Clean-Latent Prediction in Latent Diffusion Transformers
TL;DR AI
2 min readKey summary
Researchers introduced JLT, a 130M latent diffusion Transformer trained on frozen FLUX.2 VAE codes.
Under matched conditions, clean-latent prediction outperformed velocity prediction, with better geometry-aware training dynamics.
The model reached strong ImageNet 256x256 results, including FID-50K 2.50 with classifier-free guidance.
The study suggests latent diffusion training targets are not just reparameterizations—they materially affect optimization and sample quality.
