Flow-OPD: On-Policy Distillation for Flow Matching Models
TL;DR AI
2 min readKey summary
Researchers introduced Flow-OPD, a unified post-training framework for flow-matching text-to-image models.
It combines specialized teacher models, cold-start initialization, on-policy sampling, task routing, dense supervision, and manifold anchor regularization.
The method aims to better align general-purpose image generators across multiple objectives while reducing metric tradeoffs and reward hacking.
In benchmarks, Flow-OPD substantially improved performance, including on Stable Diffusion 3.5 Medium.
