Switch language한국어
Back to the list

Reinforcing Few-step Generators via Reward-Tilted Distribution Matching

TL;DR AI

Key summary

2 min read
  1. Researchers introduced RTDMD, a two-stage method for few-step image generation alignment.

  2. It first distills generators with consistency-aware distribution matching, then fine-tunes them with reward optimization.

  3. The method combines hybrid policy gradients and reduced-variance SubGRPO to better match human preferences.

  4. RTDMD achieves state-of-the-art 4-step text-to-image results while keeping inference fast across SD3, SD3.5, and FLUX.2.

Read the original