CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — agents 12 upvotes

MemBodied: Recurrent Associative Memory for Vision-Language-Action Models

QUESTION — How can long-term episodic memory be provided to Vision-Language-Action models without bloating the context or increasing inference latency?

The paper introduces MemBodied, a fixed-size episodic memory featuring two complementary components: an associative state recording interactions across policy calls and an episode anchor preserving a compact representation of the initial scene. Instead of bloating context with raw past observations, the model conditions action generation on current inputs and these memory components. Experiments across RMBench tasks and LIBERO-Long demonstrate that MemBodied substantially improves success rates over stateless policies and vanilla recurrent memories while adding very few parameters.

MemBodied achieves 7.81times the mean success rate of a stateless policy and 2.98times of vanilla recurrent memory across five evaluated RMBench tasks.

It outperforms the strongest memory-augmented baseline by 1.3times with 10times fewer added parameters.

On the fully observable LIBERO-Long suite, it reached 90.6%, a 5.4% improvement over the stateless π_0 policy.

soujanyaporia · 23 Sept 2026 read the original ↗
↑