[论文] [论文] Unified Video Dense Prediction from Disjoint Data
论文概要
研究领域: CV 作者: Yihong Sun, Seoung Wug Oh, Jiahui Huang 发布时间: 2026-07-24 arXiv: 2507.19319
中文摘要
场景理解需要同时预测几何、外观和语义。然而,现有的任务特定标注分散在不兼容的、领域特定的数据集中。当前的统一系统通过将训练限制在完全共同标注的数据上,或承担伪标注的大量计算成本来规避这一问题。为此,我们引入了UniD,一个统一视频模型,联合预测八种密集场景属性——深度、表面法线、语义分割、边界、人体部位、反照率、阴影和材质——全部从分散的、领域特定的数据集中学习。我们提出了一个简单而有效的蒸馏步骤,其中每个任务的专家通过轻量级任务投影器监督统一主干网络,消除了对标注重叠或伪标注的需求。我们的核心洞察是,预训练扩散模型的强视觉先验足以弥合分散训练源引入的领域差距,实现对训练期间从未见过的场景-任务组合的稳健泛化。UniD在各项任务中达到了与任务专家和多任务基线相当的性能,对分布外场景具有强泛化能力,并增强了时间一致性和跨任务一致性。
原文摘要
Scene understanding requires simultaneous prediction about geometry, appearance, and semantics. However, existing task-specific annotations are fragmented across incompatible, domain-specific datasets. Current unified systems circumvent this by restricting training to fully co-annotated data, or by incurring the large computational cost of pseudo-labeling. To mitigate this, we introduce UniD, a unified video model that jointly predicts eight dense scene properties-depth, surface normals, semantic segmentation, boundaries, human parts, albedo, shading, and materials-all learned from disjoint, domain-specific datasets. We propose a simple yet effective distillation step in which per-task experts supervise a unified backbone through lightweight task projectors, eliminating the need for annota...
--- *自动采集于 2026-07-25*
#论文 #arXiv #CV #小凯
🌟 智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。
🎁 领取 2000万 Tokens