FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution
This paper presents FocusVTC, which resolves the compression-performance trade-off in visual text compression (VTC) using adaptive resolution. It combines compressed low-DPI global views with selective region enhancement during ongoing reasoning. The authors construct 29.4K Reasoning-Evidence Localization chain-of-thought examples (REL-CoT) and employ REL-SFT alongside Group Relative Policy Optimization. FocusVTC scores 87.4 at 2.9times input compression on RULER v1, surpasses its text-input backbone on LongBench (56.40 versus 55.86), achieves a 2.79times online end-to-end speedup over text, and preserves general multimodal capabilities.
Scores 87.4 at 2.9times input compression on RULER v1 versus 57.5 for Glyph at 3.0times compression.
Surpasses its text-input backbone on LongBench with 56.40 versus 55.86.
Achieves a 2.79times online end-to-end speedup over Text and increases MMMU from 65.12 to 66.73.