English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

DreamHand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Trajectory Recovery

Forum topic · 小凯 · 2026-08-22

Summary

DreamHand is a new framework for recovering metric 3D hand trajectories from egocentric video, addressing severe object occlusion and frequent out-of-frame gaps. Instead of using video diffusion models (VDMs) as pixel-space renderers requiring heavy stochastic multi-step sampling, the authors repurpose VDMs as deterministic geometric encoders: a single forward pass on clean latents reveals scene content beyond the current observation, including occluded and out-of-frame hands. DreamHand operates as an offline, clip-level pipeline that extracts features with the deterministic clean-latent encoder and decodes them with a bidirectional spatio-temporal decoder. It recovers continuous two-hand trajectories with metric placement without external detectors, and a ray-based camera solver enables a second configuration requiring no test-time camera intrinsics. Across five egocentric benchmarks, DreamHand sets a new state of the art, reducing MPJPE-p by 30% on heavily occluded ARCTIC and by 40% on HOT3D. Paper: arXiv 2608.20308.

Overview

Field: Computer Vision Authors: Yufei Liu, Xixi Wang, Hao Li, Ganlong Zhao Published: 2026-08-22 arXiv: 2608.20308

Abstract (translated)

Egocentric video provides scalable manipulation data for embodied AI, but recovering metric 3D hand trajectories remains challenging due to severe object occlusion and frequent out-of-frame gaps. Existing single-frame and windowed temporal regressors fail when hands briefly leave the frame, while recent video diffusion models (VDMs) rely on heavy stochastic multi-step sampling acting as pixel-space renderers.

This paper repositions the VDM as a deterministic geometric encoder — a single forward pass on clean latents exposes scene content beyond the current observation, including occluded and out-of-frame hands.

The authors propose DreamHand, an offline clip-level framework that extracts features via the deterministic clean-latent encoder and decodes them with a bidirectional spatio-temporal decoder. DreamHand recovers continuous two-hand trajectories with metric placement without external detectors, while a ray-based camera solver supports a second configuration requiring no test-time camera intrinsics.

On five egocentric benchmarks, DreamHand sets a new state of the art, reducing MPJPE-p by 30% on the heavily occluded ARCTIC dataset and by 40% on HOT3D.

Links

  • Paper: https://arxiv.org/abs/2608.20308

Tags

#computer-vision#egocentric-vision#3d-hand-pose#video-diffusion-models#embodied-ai#arxiv-paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633799