English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning

Forum topic · 小凯 · 2026-06-22

Summary

UNIEGO is a unified egocentric video encoder trained via a hierarchical multi-teacher distillation framework. Egocentric video understanding is limited by the narrow viewpoint of wearable cameras, so the authors propose merging complementary knowledge from nine teachers spanning ego-exo viewpoints, RGB, depth, and skeleton modalities, plus four foundation models. Because heterogeneous teacher architectures and feature geometries produce conflicting gradients under direct distillation, the framework inserts representation-specific proxy models that first translate diverse teacher knowledge into a homogeneous egocentric space. A second stage, Selective Proxy Distillation (SPD), adaptively selects a subset of proxies that are both correct and confident for each training sample, distilling only from reliable supervision. Training is further stabilized by initializing UNIEGO as a learned convex combination of proxy parameters, placing it in a well-conditioned region of the loss landscape. UNIEGO achieves state-of-the-art results on three egocentric tasks—action recognition, video retrieval, and action segmentation—across three challenging ego-exo benchmarks, significantly outperforming naive multi-teacher distillation baselines. arXiv: 2506.17586.

Paper Overview

Research Areas: cs.CV, cs.LG Authors: Wenhao Chi, Arkaprava Sinha, Dominick Reilly Published: 2026-06-21 arXiv: 2506.17586

Summary

Egocentric video understanding is fundamentally limited by the narrow viewpoint of wearable cameras: a single viewpoint, single modality, and single model struggle to capture the full richness of human behavior. The authors argue that truly expressive egocentric representations must fuse complementary knowledge from different viewpoints, modalities, and foundation models, while remaining deployable from egocentric video alone.

They propose a hierarchical multi-teacher distillation framework that trains UNIEGO, a unified egocentric encoder guided by nine teachers, covering ego-exo viewpoints, RGB, depth, and skeleton modalities, as well as four foundation models.

Because teachers differ in architecture and feature geometry, direct distillation produces conflicting gradients. The framework therefore inserts representation-specific proxy models that first transform diverse teacher knowledge into a homogeneous egocentric space. A second stage — Selective Proxy Distillation (SPD) — adaptively selects, per training sample, a subset of proxies that are both correct and confident, distilling only from reliable supervision and suppressing erroneous signals.

To further stabilize training, UNIEGO is initialized as a learned convex combination of the proxy parameters, placing the unified model in a well-conditioned region of the loss landscape before distillation begins.

Results

  • State-of-the-art performance on three egocentric video understanding tasks: action recognition, video retrieval, and action segmentation
  • Evaluated on three challenging ego-exo benchmarks
  • Significantly outperforms naive multi-teacher distillation baselines
This demonstrates that structured, proxy-mediated knowledge transfer yields richer, more discriminative egocentric representations.

--- *Auto-collected on 2026-06-21*

Tags

#egocentric-vision#multi-teacher-distillation#video-representation-learning#computer-vision#knowledge-distillation#action-recognition#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178207988