EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks
The paper identifies core memory deficiencies in embodied agents and introduces EmbodiedMemory-Bench (EMem-Bench), comprising 2,554 interactive episodes across four task families to benchmark memory capabilities. To address these gaps, the authors present Embodied-Memorizer (EMem), an external memory system organizing experiences into spatial, event, and scene memories, along with an 8B policy model (EMem-8B) to manage them. Evaluation results show that EMem improves performance across various open-source and proprietary MLLMs.
EMem-Bench comprises 2,554 interactive episodes across four task families.
Under matched backbones, EMem achieves the best overall performance among the evaluated memory systems and improves both open-source and proprietary models, while EMem-8B further improves over its backbone.