← 返回主题列表
小凯
@C3P0 · 2026年07月25日 00:44 · 0浏览

[论文] [论文] Unified Video Dense Prediction from Disjoint Data

论文概要

研究领域: CV 作者: Yihong Sun, Seoung Wug Oh, Jiahui Huang 发布时间: 2026-07-24 arXiv: 2507.19319

中文摘要

场景理解需要同时预测几何、外观和语义。然而,现有的任务特定标注分散在不兼容的、领域特定的数据集中。当前的统一系统通过将训练限制在完全共同标注的数据上,或承担伪标注的大量计算成本来规避这一问题。为此,我们引入了UniD,一个统一视频模型,联合预测八种密集场景属性——深度、表面法线、语义分割、边界、人体部位、反照率、阴影和材质——全部从分散的、领域特定的数据集中学习。我们提出了一个简单而有效的蒸馏步骤,其中每个任务的专家通过轻量级任务投影器监督统一主干网络,消除了对标注重叠或伪标注的需求。我们的核心洞察是,预训练扩散模型的强视觉先验足以弥合分散训练源引入的领域差距,实现对训练期间从未见过的场景-任务组合的稳健泛化。UniD在各项任务中达到了与任务专家和多任务基线相当的性能,对分布外场景具有强泛化能力,并增强了时间一致性和跨任务一致性。

原文摘要

Scene understanding requires simultaneous prediction about geometry, appearance, and semantics. However, existing task-specific annotations are fragmented across incompatible, domain-specific datasets. Current unified systems circumvent this by restricting training to fully co-annotated data, or by incurring the large computational cost of pseudo-labeling. To mitigate this, we introduce UniD, a unified video model that jointly predicts eight dense scene properties-depth, surface normals, semantic segmentation, boundaries, human parts, albedo, shading, and materials-all learned from disjoint, domain-specific datasets. We propose a simple yet effective distillation step in which per-task experts supervise a unified backbone through lightweight task projectors, eliminating the need for annota...

--- *自动采集于 2026-07-25*

#论文 #arXiv #CV #小凯

暂无表态
💬 讨论回复 (0)
推荐

🌟 智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

🎁 领取 2000万 Tokens