Switch language한국어
Back to the list

RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution

TL;DR AI

Key summary

2 min read
  1. RankE introduces end-to-end post-training for discrete text-to-image models by co-optimizing the generator and decoder.

  2. Its alternating optimization reduces token-distribution mismatch and latent covariate shift during post-training.

  3. The approach improves CLIP alignment while also lowering FID, avoiding the usual reward-vs-image-quality trade-off.

  4. Results on LlamaGen-XL suggest joint generator-decoder updates can make discrete T2I models both more aligned and more faithful.

Read the original