English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

DreamHand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Trajectory Recovery

Forum topic · 小凯 · 2026-08-22

Summary

DreamHand is a new framework for recovering metric 3D hand trajectories from egocentric video, addressing severe object occlusion and out-of-frame gaps. Instead of using video diffusion models (VDMs) as pixel-space renderers requiring heavy stochastic multi-step sampling, the authors repurpose the VDM as a deterministic geometric encoder: a single forward pass on clean latents exposes scene content beyond the current observation, including occluded and off-frame hands. The offline segment-level framework extracts features with this deterministic clean-latent encoder and decodes them with a bidirectional spatiotemporal decoder, producing continuous two-hand trajectories with metric placement without external detectors. A ray-based camera solver enables a second configuration that requires no test-time camera intrinsics. Evaluated on five egocentric benchmarks, DreamHand sets new state-of-the-art results, reducing MPJPE-p by 30% on the heavily occluded ARCTIC dataset and 40% on HOT3D. The paper (arXiv:2608.20308) is authored by Yufei Liu, Xixi Wang, Hao Li, and Ganlong Zhao.

Overview

Field: Computer Vision Authors: Yufei Liu, Xixi Wang, Hao Li, Ganlong Zhao Published: 2026-08-22 arXiv: 2608.20308

Abstract

Egocentric video offers scalable manipulation data for embodied AI, but recovering metric 3D hand trajectories remains challenging due to severe object occlusion and frequent out-of-frame gaps. Existing single-frame and windowed temporal regressors fail when hands briefly leave the frame, while recent video diffusion models (VDMs) rely on heavy stochastic multi-step sampling as pixel-space renderers.

This paper repurposes the VDM as a deterministic geometric encoder — a single forward pass on clean latents exposes scene content beyond the current observation, including occluded and off-frame hands. The authors propose DreamHand, an offline segment-level framework that extracts features via the deterministic clean-latent encoder and decodes them with a bidirectional spatiotemporal decoder.

Key Contributions

  • Recasts video diffusion models as deterministic geometric encoders rather than stochastic pixel-space renderers
  • Recovers continuous two-hand trajectories with metric placement, without external detectors
  • A ray-based camera solver enables a second configuration requiring no test-time camera intrinsics
  • Results

    DreamHand establishes new state-of-the-art performance across five egocentric benchmarks:

  • 30% reduction in MPJPE-p on the heavily occluded ARCTIC dataset
  • 40% reduction in MPJPE-p on HOT3D
---

*Paper link: https://arxiv.org/abs/2608.20308*

Tags

#computer-vision#egocentric-vision#hand-pose-estimation#video-diffusion-models#3d-trajectory#embodied-ai#arxiv-paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633820