PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection
QUESTION — How can vision-language models achieve panoptic grounded captioning through mask proposal selection?
This work studies panoptic grounded captioning, which requires a vision-language model to describe foreground objects and background regions while grounding referring phrases with pixel-level masks. The authors introduce PanoCaps, a human-annotated benchmark built from panoptic segmentation datasets, along with a phrase-mask matching protocol and a generalized Panoptic Quality (gPQ) metric. They propose PANORAMA, a VLM that conditions a pretrained segmenter on contextualized phrase representations to select corresponding candidate masks. Joint training with caption generation enables the production of precise entity-level segmentations and detailed, mask-consistent captions.
HuggingSara · 16 Sept 2026
read the original ↗