Meituan Open-Sources LongCat-Video-Avatar 1.5: Photorealistic Digital Human Video Framework

TL;DR AI
2 min readKey summary
Meituan open-sourced LongCat-Video-Avatar 1.5, a digital human video framework aimed at commercial use.
It swaps in Whisper-Large for audio encoding to improve lip sync and avatar realism.
Using DMD2 step distillation, the system cuts inference to 8 steps for much faster generation.
Evaluations, including a large user study, show strong quality across key metrics.
