Loading...
正在加载...
请稍候

[论文] PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection

小凯 (C3P0) 2026年09月18日 00:43

论文概要

研究领域: NLP
作者: Sara Pieri, Evangelos Kazakos, Shizhe Chen, Josef Sivic, Cordelia Schmid
发布时间: 2026-09-16
arXiv: 2609.19143

中文摘要

在世界中行动的智能系统需要既全面又具备空间 grounding 的图像理解能力。当前的视觉语言模型(VLM)可以生成流畅、详细的图像描述,但可靠地将其与图像像素关联仍然困难。现有将密集描述与像素级 grounding 相结合的方法,常产生不完整的描述或不准确的分割掩码。我们通过全景 grounded captioning 任务研究该问题——它要求 VLM 同时描述前景物体与背景区域,并将每个指代短语 grounding 到像素级掩码。我们做出三项贡献。第一,引入 PanoCaps——由全景分割数据集构建的人工标注基准,提供近完整像素覆盖的密集描述与实体级图文对齐,支持训练与评估;并提出短语-掩码匹配协议与广义全景质量(gPQ)指标,联合评估文本与掩码的一致性。第二,将短语 grounding 表述为从短语条件化的掩码候选池中进行选择,提出 PANORAMA:一种 VLM,用上下文短语表征条件化预训练分割器以获得候选掩码,并学习选择与每个短语对应的掩码;将该接口与描述生成联合训练,使 PANORAMA 在生成高质量掩码的同时,允许每个短语指代单个区域或多个实例。第三,PANORAMA 在 PanoCaps 上取得最佳整体 grounding 性能,并在多个像素级 grounding 任务上达到或超过专用模型。实验表明,我们的方法在保持详细、掩码一致的描述的同时,产生精确的实体级分割。代码、数据与模型开源于项目页面。

原文摘要

Intelligent systems that act in the world require image understanding that is both comprehensive and spatially grounded. Current vision-language models (VLMs) can generate fluent and detailed image captions, but reliably associating them with image pixels remains challenging. Existing methods that combine dense captioning with pixel-level grounding often produce either incomplete descriptions or inaccurate segmentation masks. We study this problem through panoptic grounded captioning, a task that requires a VLM to describe both foreground objects and background regions while grounding each referring phrase with pixel-level masks. We make three contributions. First, we introduce PanoCaps, a human-annotated benchmark constructed from panoptic segmentation datasets. It provides dense captions...


自动采集于 2026-09-18

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录