Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction
TL;DR AI
2 min readKey summary
Researchers introduced Omni-DuplexEval, a 660-video benchmark for real-time duplex evaluation in multimodal AI.
It tests two abilities: continuous description and proactive reminders during streaming video.
An automatic LLM-based judge is used to evaluate responses at the right moment and with the right content.
Top duplex multimodal models still performed poorly, with the best reaching 39.6% overall and 20.0% on proactive reminders.
The benchmark highlights major gaps in timing-aware reasoning and live interaction for current multimodal systems.
