CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — multimodal 32 upvotes

Region-Level Policy Optimization for Fine-grained MLLM Perception

QUESTION — How can a proposal network be optimized with region-level reinforcement learning to improve fine-grained visual perception in MLLMs without inflating token costs?

This research shows that region of interest (RoI) localization and content recognition in MLLMs have different resolution requirements, with localization tolerating roughly 3 to 4 times stronger token compression. To optimize the proposal network without region annotations, the authors introduce Vision-RL2, a region-level reinforcement learning method treating coherent regions as actions. A frozen MLLM reader scores each region by how its removal changes the answer likelihood. This refines the predictor to enable sparse encoding while excluding background tokens. Across six fine-grained benchmarks and four MLLM backbones, Vision-RL2 improves accuracy while using about 4 times fewer visual tokens.

Vision-RL2 improves accuracy over the base model at every token budget and surpasses its largest-budget accuracy with about 4 times fewer visual tokens.

YuhengSSS · 17 Sept 2026 read the original ↗
↑