Switch language한국어
Back to the list

MEME: Multi-entity & Evolving Memory Evaluation

TL;DR AI

Key summary

2 min read
  1. Researchers introduced MEME, a six-task benchmark for evaluating LLM agent memory in multi-entity, evolving scenarios.

  2. The benchmark adds new dependency, deletion, and other state-changing tasks to test realistic memory reasoning beyond simple retrieval.

  3. Across six memory systems and 100 controlled episodes, all systems struggled badly with dependency reasoning under default settings.

  4. Only a file-based agent using Claude Opus 4.7 improved results somewhat, and only at very high cost.

  5. The findings show a clear gap between strong-looking retrieval performance and robust memory for real-world, changing contexts.

Read the original