Switch language한국어
Back to the list

SEGA: Spectral-Energy Guided Attention for Resolution Extrapolation in Diffusion Transformers

TL;DR AI

Key summary

2 min read
  1. SEGA is a training-free method for high-resolution text-to-image generation with diffusion transformers.

  2. It uses the latent’s spatial-frequency structure to adaptively scale attention across RoPE components during denoising.

  3. This helps diffusion transformers extrapolate beyond their training resolution range while preserving global structure and fine detail.

  4. The result is better high-resolution synthesis than prior training-free baselines, without extra training.

Read the original