English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

LIGHT: Classifier-Free Guidance from Denoising Pace for Human-Object Interaction Animation

Forum topic · 小凯 · 2026-03-28

Summary

Generating realistic human-object interaction (HOI) animations remains challenging because it requires jointly modeling dynamic human actions and diverse object geometries. Prior diffusion-based approaches often rely on hand-crafted contact priors or human-imposed kinematic constraints to improve contact quality. LIGHT, proposed by researchers including Ziyin Wang, Sirui Xu, Chuan Guo, and Liang-Yan Gui (arXiv:2603.25734), offers a data-driven alternative in which guidance emerges from the denoising pace itself, reducing dependence on manually designed priors. Built on diffusion forcing, LIGHT factors the representation into modality-specific components and assigns individualized noise levels with asynchronous denoising schedules. Cleaner components guide noisier ones through cross-attention, yielding guidance without auxiliary classifiers. The authors show this data-driven guidance is inherently contact-aware and improves further when training data is augmented with diverse synthetic object geometries, encouraging invariance of contact semantics to geometric variation. Experiments demonstrate that this pace-induced guidance captures contact priors more effectively than conventional classifier-free guidance, achieving higher contact fidelity, more realistic HOI generation, and stronger generalization to unseen objects and tasks.

Overview

Field: Computer Vision Authors: Ziyin Wang, Sirui Xu, Chuan Guo, Bing Zhou, Jiangshan Gong, Jian Wang, Yu-Xiong Wang, Liang-Yan Gui Published: 2026-03-26 arXiv: 2603.25734v1

Key Points

  • Problem: Generating realistic human-object interaction (HOI) animations requires jointly modeling dynamic human actions and diverse object geometries; prior diffusion-based methods rely on hand-crafted contact priors or human-imposed kinematic constraints.
  • Proposed method — LIGHT: A data-driven alternative where guidance emerges from the denoising pace itself, reducing dependence on manually designed priors.
  • Mechanism: Building on diffusion forcing, LIGHT factors the representation into modality-specific components with individualized noise levels and asynchronous denoising schedules. Cleaner components guide noisier ones through cross-attention, yielding guidance without auxiliary classifiers.
  • Contact awareness: This data-driven guidance is inherently contact-aware, and is further enhanced when training data is augmented with a broad range of synthetic object geometries, encouraging invariance of contact semantics to geometric diversity.
  • Results: Extensive experiments show that pace-induced guidance reflects contact prior advantages more effectively than conventional classifier-free guidance, achieving higher contact fidelity, more realistic HOI generation, and stronger generalization to unseen objects and tasks.

Original Abstract (excerpt)

> Generating realistic human-object interaction (HOI) animations remains challenging because it requires jointly modeling dynamic human actions and diverse object geometries. Prior diffusion-based approaches often rely on hand-crafted contact priors or human-imposed kinematic constraints to improve contact quality. We propose LIGHT, a data-driven alternative in which guidance emerges from the denoising pace itself, reducing dependence on manually designed priors. Building on diffusion forcing, we factor the representation into modality-specific components and assign individualized noise levels with asynchronous denoising schedules. In this paradigm, cleaner components guide noisier ones through cross-attention, yielding guidance without auxiliary classifiers...

---

*Auto-collected on 2026-03-28*

Tags

#computer-vision#diffusion-models#human-object-interaction#motion-generation#animation#guidance#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169366