English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

UniD: Unified Video Dense Prediction from Disjoint Data

Forum topic · 小凯 · 2026-07-25

Summary

UniD is a unified video model that jointly predicts eight dense scene properties—depth, surface normals, semantic segmentation, boundaries, human parts, albedo, shading, and materials—trained from disjoint, domain-specific datasets. Existing unified systems typically require fully co-annotated data or costly pseudo-labeling; UniD avoids both by using a simple distillation step in which per-task experts supervise a unified backbone through lightweight task projectors, eliminating the need for annotation overlap or pseudo-labels. The key insight is that strong visual priors from a pretrained diffusion model are sufficient to bridge the domain gap introduced by disjoint training sources, enabling robust generalization to scene-task combinations never seen during training. UniD achieves performance comparable to task-specific experts and multi-task baselines across tasks, generalizes well to out-of-distribution scenes, and improves temporal and cross-task consistency. Paper: arXiv 2507.19319.

论文概要

Research Area: Computer Vision Authors: Yihong Sun, Seoung Wug Oh, Jiahui Huang Published: 2026-07-24 arXiv: 2507.19319

Introduction

Scene understanding requires simultaneous prediction of geometry, appearance, and semantics. However, existing task-specific annotations are fragmented across incompatible, domain-specific datasets. Current unified systems work around this by restricting training to fully co-annotated data, or by incurring the large computational cost of pseudo-labeling.

Method

The authors introduce UniD, a unified video model that jointly predicts eight dense scene properties:

  • Depth
  • Surface normals
  • Semantic segmentation
  • Boundaries
  • Human parts
  • Albedo
  • Shading
  • Materials
  • All of these are learned from disjoint, domain-specific datasets. The core of the method is a simple yet effective distillation step: per-task experts supervise a unified backbone through lightweight task projectors, eliminating the need for annotation overlap or pseudo-labels.

    Key Insight

    The strong visual priors of a pretrained diffusion model are sufficient to bridge the domain gap introduced by disjoint training sources, enabling robust generalization to scene-task combinations never seen during training.

    Results

  • Matches task-specific experts and multi-task baselines across all eight tasks
  • Strong generalization to out-of-distribution scenes
  • Improved temporal consistency and cross-task consistency
  • Links

  • arXiv: 2507.19319
--- *Auto-collected on 2026-07-25.*

Tags

#computer-vision#diffusion-models#dense-prediction#multi-task-learning#knowledge-distillation#video-understanding#unid#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178447082