English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

UNIEGO: Proxies as Mediators for Unified Egocentric Video Representations

Forum topic · 小凯 · 2026-06-20

Summary

UNIEGO (arXiv:2506.16806) is a unified egocentric video encoder trained via a hierarchical multi-teacher distillation framework. Egocentric video understanding is limited by the narrow field of wearable cameras, so the authors distill knowledge from nine teachers covering ego-exo viewpoints, RGB, depth, and skeleton modalities, plus four foundation models. Instead of distilling directly from heterogeneous teachers—whose incompatible architectures and feature geometries cause gradient conflicts—the framework inserts modality-specific proxy models that translate diverse teacher knowledge into a homogeneous egocentric space. A second stage, Selective Proxy Distillation (SPD), adaptively selects correct and confident proxies per training sample to provide reliable supervision and suppress erroneous signals, while initializing UNIEGO as a learned convex combination of proxy parameters to keep training in a well-conditioned loss landscape. UNIEGO achieves state-of-the-art results on three egocentric tasks—action recognition, video retrieval, and action segmentation—across three challenging ego-exo benchmarks, outperforming naive multi-teacher distillation baselines.

Paper Overview

Field: Computer Vision Authors: Wenhao Chi, Arkaprava Sinha, Dominick Reilly Published: 2025-06-20 arXiv: 2506.16806

Summary

Egocentric video understanding is inherently limited by the narrow perspective of wearable cameras: a single viewpoint, a single modality, or a single model cannot capture the full richness of human action. The authors argue that a truly expressive egocentric representation must subsume complementary knowledge across viewpoints, modalities, and foundation model representations, while remaining deployable from egocentric video alone.

To this end, they introduce a hierarchical multi-teacher distillation framework that produces UNIEGO, a unified egocentric encoder trained with nine teachers spanning ego-exo viewpoints, RGB, depth, and skeleton modalities, and four foundation models.

Rather than distilling directly from heterogeneous teachers—whose incompatible architectures and feature geometries induce conflicting gradients—the framework inserts representation-specific proxy models that translate diverse teacher knowledge into a homogeneous egocentric space.

In a second distillation stage, Selective Proxy Distillation (SPD) adaptively selects a subset of proxies that are both correct and confident for each training sample, so that only reliable supervision drives the distillation and erroneous signals are suppressed. SPD is further stabilized by initializing UNIEGO as a learned convex combination of the proxy parameters, placing the unified model in a well-conditioned region of the loss landscape before distillation begins.

Results

UNIEGO achieves state-of-the-art performance on three egocentric video understanding tasks—action recognition, video retrieval, and action segmentation—across three challenging ego-exo benchmarks, outperforming naive multi-teacher distillation baselines. This confirms that structured, proxy-mediated knowledge transfer yields richer and more discriminative egocentric representations.

*Auto-collected on 2026-06-20.*

Tags

#egocentric-video#knowledge-distillation#multi-teacher-learning#computer-vision#video-understanding#representation-learning#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981549