SafeDiffusion-R1: Online Reward Steering for Safe Diffusion Post-Training

TL;DR AI
2 min readKey summary
Researchers introduced SafeDiffusion-R1, an online RL post-training method for diffusion models that uses GRPO and a CLIP-based steering reward.
It learns from positive and negative prompts to reduce unsafe and nudity-related generations across both in-domain and out-of-domain harms.
The approach improves compositional image quality while avoiding paired supervised safety data and specialized reward-model tuning.
Overall, it offers a scalable alternative to data-heavy safety fine-tuning and helps limit catastrophic forgetting in image generation models.
