Switch language한국어
Back to the list

Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization

TL;DR AI

Key summary

2 min read
  1. Researchers introduced BiDPO, a preference-based fine-tuning framework for text-to-image models.

  2. It builds BiComp, a quality-controlled preference dataset, and extends Diffusion DPO to optimize both image and text preferences.

  3. BiDPO also adds region-level guidance to better handle compositional concepts such as object relations, attributes, and counting.

  4. Experiments show stronger compositional fidelity across benchmarks than prior methods.

Read the original