EVA01: Unified Native 3D Understanding and Generation via Mixture-of-Transformers
TL;DR AI
2 min readKey summary
EVA01 is a unified multimodal framework that natively brings 3D mesh understanding, generation, and editing into language models.
It uses separate understanding and generation experts with shared attention and hard routing to better align semantic and geometric representations.
The system reports state-of-the-art results in native text-to-3D generation and multi-turn geometric editing.
The work suggests a path to make 3D meshes a first-class part of multimodal models, reducing reliance on separate reconstruction pipelines.
