MEME: Multi-entity & Evolving Memory Evaluation

TL;DR AI
2 min readKey summary
Researchers introduced MEME, a six-task benchmark for evaluating LLM agent memory in multi-entity, evolving scenarios.
The benchmark adds new dependency, deletion, and other state-changing tasks to test realistic memory reasoning beyond simple retrieval.
Across six memory systems and 100 controlled episodes, all systems struggled badly with dependency reasoning under default settings.
Only a file-based agent using Claude Opus 4.7 improved results somewhat, and only at very high cost.
The findings show a clear gap between strong-looking retrieval performance and robust memory for real-world, changing contexts.
