CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — multimodal 21 upvotes

FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution

QUESTION — How can visual text compression with adaptive resolution reduce tokens while preserving reasoning capabilities?

This paper presents FocusVTC, which resolves the compression-performance trade-off in visual text compression (VTC) using adaptive resolution. It combines compressed low-DPI global views with selective region enhancement during ongoing reasoning. The authors construct 29.4K Reasoning-Evidence Localization chain-of-thought examples (REL-CoT) and employ REL-SFT alongside Group Relative Policy Optimization. FocusVTC scores 87.4 at 2.9times input compression on RULER v1, surpasses its text-input backbone on LongBench (56.40 versus 55.86), achieves a 2.79times online end-to-end speedup over text, and preserves general multimodal capabilities.

Scores 87.4 at 2.9times input compression on RULER v1 versus 57.5 for Glyph at 3.0times compression.

Surpasses its text-input backbone on LongBench with 56.40 versus 55.86.

Achieves a 2.79times online end-to-end speedup over Text and increases MMMU from 65.12 to 66.73.

zfz04 · 29 Sept 2026 read the original ↗
↑