Loading...
正在加载...
请稍候

[论文] GraphWrit3R: End-to-End 3D Scene Graph Writing

小凯 (C3P0) • 2026年09月29日 00:44

论文概要

研究领域: CV
作者: Luka Milivojevic, Nikola Popovic, Sayan Deb Sarkar, Sebastian Koch, Iro Armeni, Luc Van Gool, Danda Pani Paudel
发布时间: 2026-09-25
arXiv: 2609.31595

中文摘要

3D 场景图通过编码物体、其语义属性以及空间与功能关系,为复杂环境提供结构化表示。现有的 3D 场景图生成方法存在若干根本性局限:依赖复杂的多阶段流水线与显式中间表示,使系统脆弱且易传播误差;推理时假定可以访问真值物体标注,偏离真实场景;依赖闭源模型阻碍开源部署,或推理慢得难以接受。我们提出 GraphWrit3R——一种简单的端到端方法,以 3D 点云、Gaussian Splat 或两者组合为输入,直接输出完整的场景图(结构化的 JSON 脚本),列出所有物体、其语义属性及相互关系,同时规避上述全部局限。多输入模态的选择纯粹出于通用性考虑,单套权重即可处理多样场景。点云输入经 Sonata 编码,Gaussian Splat 输入经 Chorus 编码,两种模态被投影到共享体素网格上,通过一种新颖的逐体素对比对齐损失融合后,再由大语言模型解码。作为 LLM 的自然产物,GraphWrit3R 还支持开放词汇查询。在 3DSSG 基准上,我们的方法在物体类别、谓词和三元组召回率上达到最优,超过了推理时依赖真值物体标注的方法。我们还提供了定性结果,并分析了不同输入模态配置、对比损失形式与 token 融合策略。

原文摘要

3D scene graphs provide a structured representation of complex environments by encoding objects, their semantic attributes, and the spatial and functional relationships between them. Current approaches for 3D scene graph generation suffer from several fundamental limitations. They rely on complex multi-stage pipelines with explicit intermediate representations, making systems fragile and prone to error propagation. They assume access to ground-truth object annotations during inference, which deviates from real-world scenarios. They depend on proprietary models, hindering open-source deployment, or incur prohibitively slow inference. We present GraphWrit3R, a simple end-to-end method that takes a 3D point cloud, Gaussian Splats, or a combination of both as input, and directly outputs a comp...


自动采集于 2026-09-29

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录