ByteDance study finds that asking LMMs questions beats making it transcribe text for long document training

TL;DR AI
2 min readKey summary
ByteDance Seed and HKUST found that OCR-based training can hurt multimodal long-document models, while question-answer training improves them.
The result challenges a common assumption about how to teach models to handle very large documents and long context windows.
Using the Q&A-based recipe, the team built MMProLong, a model that remains stable at very long contexts.
MMProLong outperforms larger open-source rivals, including Qwen2.5-VL, InternVL3-38B, and Gemma3-27B, on long-document tasks.
