Switch language한국어
Back to the list

ByteDance study finds that asking LMMs questions beats making it transcribe text for long document training

TL;DR AI

Key summary

2 min read
  1. ByteDance Seed and HKUST found that OCR-based training can hurt multimodal long-document models, while question-answer training improves them.

  2. The result challenges a common assumption about how to teach models to handle very large documents and long context windows.

  3. Using the Q&A-based recipe, the team built MMProLong, a model that remains stable at very long contexts.

  4. MMProLong outperforms larger open-source rivals, including Qwen2.5-VL, InternVL3-38B, and Gemma3-27B, on long-document tasks.

Read the original