Reinforcing Few-step Generators via Reward-Tilted Distribution Matching

TL;DR AI
2 min readKey summary
Researchers introduced RTDMD, a two-stage method for aligning few-step text-to-image generators with human preferences.
It first improves generator tracking via ambient-consistent distribution matching, then applies a hybrid policy-gradient reward optimization stage.
On SD3, SD3.5, and FLUX.2, the method achieved state-of-the-art results for four-step image generation.
The gains were especially strong on standard aesthetic and compositional benchmarks, while keeping inference fast.
