TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward
TL;DR AI
2 min readKey summary
Researchers introduced TILT, a training-free method for compositional text-to-image generation with diffusion models.
TILT adjusts sampling at test time using a model-intrinsic reward, aiming to handle prompts with multiple concepts that are not jointly represented.
It derives a KL-constrained tilted target distribution and combines two guidance strategies into a hybrid approach.
On T2ICompBench, TILT beat prior baselines on compositional alignment while preserving visual quality.
The method offers a practical way to improve prompt following without extra supervision or external reward models.
