Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models
The authors investigate optimizing multimodal large language models (MLLMs) by performing low-rank interventions on visual-to-text information flow. Observing that relevant visual influence is concentrated in a low-dimensional subspace and layer-specific visual states can be approximated by lightweight MLPs, they propose δ-Vision. This method replaces repeated Transformer evolution of visual tokens with lightweight low-rank adapters that construct layer-wise visual memories while preserving all visual tokens for text retrieval.
Restoring only a few directions after visual-to-text attention is blocked recovers most of the lost accuracy.
Lightweight MLPs approximate layer-specific visual states with high cosine similarity and low reconstruction error.
δ-Vision achieves higher accuracy than visual token pruning baselines at comparable or lower computation.