CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 14 upvotes

The Low-Rank Structure of VLA Reinforcement Learning

QUESTION — What parameter structures within vision-language-action models are responsible for performance gains during reinforcement learning post-training?

This study investigates how reinforcement learning (RL) reshapes policies in vision-language-action (VLA) models. The authors discover that RL induces low-rank parameter updates highly concentrated in the action expert's Timestep Modules across models such as $\pi_{0.5}$ and GR00T N1.5/N1.6. Through systematic module-replacement experiments, they show these modules capture a disproportionate share of performance gains from RL. Furthermore, shift update directions within these modules strongly predict task success with ROC-AUC up to 99.6%, enabling further policy improvements without additional RL training.

RL induces low-rank parameter updates concentrated in the action expert's Timestep Modules.

Shift update directions strongly predict task success with ROC-AUC up to 99.6%.

Jongwondd · 28 Sept 2026 read the original ↗
↑