PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models
The paper introduces PhysBrain 1.5, a unified model designed to understand physical environments, generate actions, and predict future states. Built on top of a general vision-language model, it encodes language, end-effector motion, and dense visual targets into discrete sequences optimized via autoregressive next-token prediction. Pre-training relies entirely on human interaction videos, followed by supervised fine-tuning on demonstrations, robot trajectories, and simulated data. Achieving an average score of 72.5 across 28 embodied benchmarks, the 8B model sets a new open-source state of the art.
Our 8B model achieves an average score of 72.5 across 28 embodied understanding benchmarks.
It achieves the best open-source results on 14 benchmarks while retaining general multimodal capabilities.
It produces end-effector trajectories and predicts future scenes through spatially aligned RGB, depth, and robot-mask outputs.