CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — multimodal 17 upvotes

PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection

QUESTION — How can vision-language models achieve panoptic grounded captioning through mask proposal selection?

This work studies panoptic grounded captioning, which requires a vision-language model to describe foreground objects and background regions while grounding referring phrases with pixel-level masks. The authors introduce PanoCaps, a human-annotated benchmark built from panoptic segmentation datasets, along with a phrase-mask matching protocol and a generalized Panoptic Quality (gPQ) metric. They propose PANORAMA, a VLM that conditions a pretrained segmenter on contextualized phrase representations to select corresponding candidate masks. Joint training with caption generation enables the production of precise entity-level segmentations and detailed, mask-consistent captions.

HuggingSara · 16 Sept 2026 read the original ↗
↑