RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution
TL;DR AI
2 min readKey summary
RankE introduces end-to-end post-training for discrete text-to-image models by co-optimizing the generator and decoder.
Its alternating optimization reduces token-distribution mismatch and latent covariate shift during post-training.
The approach improves CLIP alignment while also lowering FID, avoiding the usual reward-vs-image-quality trade-off.
Results on LlamaGen-XL suggest joint generator-decoder updates can make discrete T2I models both more aligned and more faithful.
