Memory is one of the most important problems in robotics. Long horizon memory is key for a variety of robot manipulation problems. However, there exist no good benchmarks for understanding progress in how well generalist robot policies can understand language.
Yinpei Dai and Yuejiang Liu made RoboMME as a solution: it’s a large benchmark which shows 16 different robot tasks, like counting objects or mastering timing. They show 14 different memory-augmented generalist policies across these different benchmarks. It’s an incredibly thorough and interesting result, aimed at driving forward this core robotic capability.
To learn more, watch Episode 91 of RoboPapers with Michael Cho and Chris Paxton!
Abstract
Memory is critical for long-horizon and history-dependent robotic manipulation. Such tasks often involve counting repeated actions or manipulating objects that become temporarily occluded. Recent vision-language-action (VLA) models have begun to incorporate memory mechanisms; however, their evaluations remain confined to narrow, non-standardized settings. This limits their systematic understanding, comparison, and progress measurement. To address these challenges, we introduce RoboMME: a large-scale standardized benchmark for evaluating and advancing VLA models in long-horizon, history-dependent scenarios. Our benchmark comprises 16 manipulation tasks constructed under a carefully designed taxonomy that evaluates temporal, spatial, object, and procedural memory. We further develop a suite of 14 memory-augmented VLA variants built on the π0.5 backbone to systematically explore different memory representations across multiple integration strategies. Experimental results show that the effectiveness of memory representations is highly task-dependent, with each design offering distinct advantages and limitations across different tasks.









