The Low-Rank Structure of VLA Reinforcement Learning
This study investigates how reinforcement learning (RL) reshapes policies in vision-language-action (VLA) models. The authors discover that RL induces low-rank parameter updates highly concentrated in the action expert's Timestep Modules across models such as $\pi_{0.5}$ and GR00T N1.5/N1.6. Through systematic module-replacement experiments, they show these modules capture a disproportionate share of performance gains from RL. Furthermore, shift update directions within these modules strongly predict task success with ROC-AUC up to 99.6%, enabling further policy improvements without additional RL training.
RL induces low-rank parameter updates concentrated in the action expert's Timestep Modules.
Shift update directions strongly predict task success with ROC-AUC up to 99.6%.