Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence
QUESTION — How can supervision of latent reasoning trajectories in MLLMs be improved?
The paper investigates latent-token behavior in MLLMs and identifies a latent evidence-credit gap, where latent tokens respond weakly to image perturbations. To bridge this gap, the authors propose ReaLVR, which introduces visual-evidence supervision into the model's free-running latent trajectories by contrasting correct and model-generated wrong answers. This approach achieves top performance on Qwen2.5-VL-7B and scales effectively to frontier models up to 235B parameters.
Achieving the highest five-task average of 63.7% on Qwen2.5-VL-7B.
Scaling visual reasoning in latent space to frontier model scales up to 235B.
MarkShaw99 · 28 Sept 2026
read the original ↗