Summary
UniEGO is a unified egocentric video encoder introduced in arXiv paper 2506.18497 (June 23, 2025) by Wenhao Chi, Arkaprava Sinha, and Dominick Reilly. Egocentric video understanding is limited by the narrow perspective of wearable cameras—a single viewpoint, modality, or model cannot capture the full richness of human action. The authors propose a hierarchical multi-teacher distillation framework that trains UniEGO with nine teachers spanning ego- and third-person viewpoints, RGB, depth, and skeleton modalities, plus four foundation models. Instead of distilling directly from heterogeneous teachers, the framework inserts representation-specific agent models that translate diverse teacher knowledge into a homogeneous egocentric space. A second-stage Selective Proxy Distillation (SPD) adaptively selects a subset of agents that are both correct and confident for each training sample, distilling only from reliable supervision and suppressing erroneous signals. UniEGO achieves state-of-the-art performance on action recognition, video retrieval, and action segmentation across three challenging ego-exo benchmarks while remaining deployable from egocentric video alone.
Paper Overview
Field: Computer Vision
Authors: Wenhao Chi, Arkaprava Sinha, Dominick Reilly
Published: 2025-06-23
arXiv: 2506.18497
Abstract
Egocentric video understanding is inherently limited by the narrow perspective of wearable cameras: a single viewpoint, a single modality, a single model cannot capture the full richness of human action. The authors argue that a truly expressive egocentric representation must subsume complementary knowledge across viewpoints, modalities, and foundation model representations, yet remain deployable from egocentric video alone.
To this end, the paper introduces a hierarchical multi-teacher distillation framework that produces UniEGO, a unified egocentric encoder trained with nine teachers spanning:
- Ego- and third-person viewpoints
- RGB, depth, and skeleton modalities
- Four foundation models
Rather than distilling directly from heterogeneous teachers—whose incompatible architectures and feature geometries would induce conflicting gradients—the framework inserts representation-specific
agent models that translate diverse teacher knowledge into a homogeneous egocentric space.
In a second stage, Selective Proxy Distillation (SPD) adaptively selects a subset of agents that are both correct and confident for each training sample, distilling only from reliable supervision and suppressing erroneous signals.
Results
UniEGO achieves state-of-the-art performance on action recognition, video retrieval, and action segmentation across three challenging ego-exo benchmarks.
---
*Auto-collected on 2026-06-23*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178208030