CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — architecture 12 upvotes

Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models

QUESTION — Can all visual tokens be preserved in MLLMs while reducing the computational cost of repeatedly evolving visual representations through the Transformer?

The authors investigate optimizing multimodal large language models (MLLMs) by performing low-rank interventions on visual-to-text information flow. Observing that relevant visual influence is concentrated in a low-dimensional subspace and layer-specific visual states can be approximated by lightweight MLPs, they propose δ-Vision. This method replaces repeated Transformer evolution of visual tokens with lightweight low-rank adapters that construct layer-wise visual memories while preserving all visual tokens for text retrieval.

Restoring only a few directions after visual-to-text attention is blocked recovers most of the lost accuracy.

Lightweight MLPs approximate layer-specific visual states with high cosine similarity and low reconstruction error.

δ-Vision achieves higher accuracy than visual token pruning baselines at comparable or lower computation.

huaXiaKyrie · 28 Sept 2026 read the original ↗
↑