Stitched Value Model for Diffusion Alignment
TL;DR AI
2 min readKey summary
Researchers introduced StitchVM, a stitching framework that adapts a truncated pixel-space reward model to a frozen diffusion backbone so it can score noisy latents.
The method uses a lightweight bridge, including Tweedie-style approximation and Monte Carlo estimation, to reuse strong pretrained reward models in latent space.
StitchVM adds only a small training cost while making alignment methods like DPS and DiffusionNFT faster and more memory efficient.
By moving reward evaluation into the latent regime, it lowers the cost and bias of diffusion model alignment and supports stronger steering.
