CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — architecture 208 upvotes

Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence

QUESTION — How can supervision of latent reasoning trajectories in MLLMs be improved?

The paper investigates latent-token behavior in MLLMs and identifies a latent evidence-credit gap, where latent tokens respond weakly to image perturbations. To bridge this gap, the authors propose ReaLVR, which introduces visual-evidence supervision into the model's free-running latent trajectories by contrasting correct and model-generated wrong answers. This approach achieves top performance on Qwen2.5-VL-7B and scales effectively to frontier models up to 235B parameters.

Achieving the highest five-task average of 63.7% on Qwen2.5-VL-7B.

Scaling visual reasoning in latent space to frontier model scales up to 235B.

MarkShaw99 · 28 Sept 2026 read the original ↗
↑