MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models

TL;DR AI
2 min readKey summary
Researchers introduced MEMLENS, a 789-question benchmark for testing long-term multimodal memory in large vision-language models.
The benchmark covers five memory abilities across four context lengths in multi-session conversations.
They evaluated 27 LVLMs and 7 memory-augmented agents, and found that neither long-context models nor memory agents alone perform well.
The results suggest future systems will need both long-context attention and structured multimodal retrieval to handle this task effectively.
