Reinforcing Few-step Generators via Reward-Tilted Distribution Matching
TL;DR AI
2 min readKey summary
Researchers introduced RTDMD, a two-stage method for few-step image generation alignment.
It first distills generators with consistency-aware distribution matching, then fine-tunes them with reward optimization.
The method combines hybrid policy gradients and reduced-variance SubGRPO to better match human preferences.
RTDMD achieves state-of-the-art 4-step text-to-image results while keeping inference fast across SD3, SD3.5, and FLUX.2.
