Region-Level Policy Optimization for Fine-grained MLLM Perception
This research shows that region of interest (RoI) localization and content recognition in MLLMs have different resolution requirements, with localization tolerating roughly 3 to 4 times stronger token compression. To optimize the proposal network without region annotations, the authors introduce Vision-RL2, a region-level reinforcement learning method treating coherent regions as actions. A frozen MLLM reader scores each region by how its removal changes the answer likelihood. This refines the predictor to enable sparse encoding while excluding background tokens. Across six fine-grained benchmarks and four MLLM backbones, Vision-RL2 improves accuracy while using about 4 times fewer visual tokens.
Vision-RL2 improves accuracy over the base model at every token budget and surpasses its largest-budget accuracy with about 4 times fewer visual tokens.