Paper Overview
- Field: Computer Vision (CV)
- Authors: Liang Xu, Chengqun Yang, Zili Lin, Xintao Lv
- Published: 2026-08-22
- arXiv: 2608.20312
- Hierarchical fine-grained textual descriptions
- Interaction categories
- Causal interaction ordering
- Subject relationships and personalities
- Vertex-level contact maps
- Physics-regularized constraints
Abstract
The ability to perceive and synthesize human-human interaction is fundamental to developing intelligent digital human systems. However, existing datasets and modeling approaches are severely limited by low-fidelity kinematics, a lack of dexterous hand gestures, and insufficient rich multimodal annotation.
This paper presents Inter-X++, a comprehensive large-scale benchmark captured with a novel hybrid motion capture system. It provides 11,388 high-fidelity interaction sequences and over 8.1 million frames, featuring precise full-body motion and detailed finger articulation.
The dataset is enriched with multi-faceted annotations, including:
Contributions
1. Unified testbed: Using these annotations, the authors define four families of downstream tasks, symmetrically spanning both generation and perception paradigms. 2. OpenHHI framework: A unified HHI representation and modeling framework that jointly optimizes interaction reconstruction and semantic understanding. 3. State-of-the-art results: Extensive experiments show that OpenHHI achieves state-of-the-art performance across both generation and perception tasks.
---
*Source link: https://arxiv.org/abs/2608.20312*