Lance: Unified Multimodal Modeling by Multi-Task Synergy

TL;DR AI
2 min readKey summary
Researchers introduced Lance, a lightweight unified multimodal model that handles image and video understanding, generation, and editing through multi-task training.
Lance was trained from scratch with a dual-stream mixture-of-experts architecture, staged multi-task learning, and modality-aware positional encoding.
The paper reports stronger open-source performance on image and video generation while preserving solid multimodal understanding.
It offers a practical path toward unifying understanding and generation across images and videos without relying solely on larger models.
