CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — multimodal 170 upvotes

PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models

QUESTION — How can a vision-language model be adapted into a unified physical foundation model for environment understanding, action generation, and future state prediction?

The paper introduces PhysBrain 1.5, a unified model designed to understand physical environments, generate actions, and predict future states. Built on top of a general vision-language model, it encodes language, end-effector motion, and dense visual targets into discrete sequences optimized via autoregressive next-token prediction. Pre-training relies entirely on human interaction videos, followed by supervised fine-tuning on demonstrations, robot trajectories, and simulated data. Achieving an average score of 72.5 across 28 embodied benchmarks, the 8B model sets a new open-source state of the art.

Our 8B model achieves an average score of 72.5 across 28 embodied understanding benchmarks.

It achieves the best open-source results on 14 benchmarks while retaining general multimodal capabilities.

It produces end-effector trajectories and predicts future scenes through spatially aligned RGB, depth, and robot-mask outputs.

VLyb · 14 Sept 2026 read the original ↗
↑