CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 11 upvotes

Think Like a World Model, Act Like a VLA: Distilling World-Model Representations into Compact Robot Policies

QUESTION — How can physical world representations be distilled from a world model into compact Vision-Language-Action policies without adding runtime inference latency?

This work proposes distilling internal representations from a frozen world model into compact Vision-Language-Action (VLA) policies by adding a single feature-alignment term during training. Once trained, the projection components are discarded, leaving the deployed policy identical in architecture to the undistilled baseline while inheriting physical grounding. Running at 32 ms on an RTX 5090, a 0.8B student model achieves high performance on LIBERO and improves manipulation tasks on RoboCasa-GR1.

The deployed policy runs in 32 ms and 1.86 GB on a consumer RTX 5090, identical to the undistilled baseline.

A 0.8B student reaches 97.9% on LIBERO.

It improves from 48.2% to 50.5% on RoboCasa-GR1 humanoid manipulation.

termanteus · 21 Sept 2026 read the original ↗
↑