Switch language한국어
Back to the list

TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward

TL;DR AI

Key summary

2 min read
  1. Researchers introduced TILT, a training-free method for compositional text-to-image generation with diffusion models.

  2. TILT adjusts sampling at test time using a model-intrinsic reward, aiming to handle prompts with multiple concepts that are not jointly represented.

  3. It derives a KL-constrained tilted target distribution and combines two guidance strategies into a hybrid approach.

  4. On T2ICompBench, TILT beat prior baselines on compositional alignment while preserving visual quality.

  5. The method offers a practical way to improve prompt following without extra supervision or external reward models.

Read the original