Summary
Inter-X++ is a comprehensive large-scale benchmark for perceiving and synthesizing human-human interaction (HHI), collected via a novel hybrid motion capture system. It provides 11,388 high-fidelity interaction sequences with over 8.1 million frames, featuring accurate whole-body motion and detailed finger articulation. The dataset is enriched with multifaceted annotations, including hierarchical fine-grained textual descriptions, interaction categories, causal interaction order, subject relationships and personalities, vertex-level contact maps, and physics-based regularization constraints. The authors establish a unified testbed covering four categories of downstream tasks spanning both generation and perception paradigms, and introduce OpenHHI, a unified HHI representation and modeling framework that jointly optimizes interaction reconstruction and semantic understanding. Extensive experiments show OpenHHI achieves state-of-the-art performance across both generative and perceptual tasks. Paper: arXiv 2608.20312.
Paper Overview
Field: Computer Vision (CV)
Authors: Liang Xu, Chengqun Yang, Zili Lin, Xintao Lv
Published: 2026-08-22
arXiv: 2608.20312
Abstract
The ability to perceive and synthesize human-human interaction (HHI) is fundamental to developing intelligent digital human systems. However, existing datasets and modeling methods suffer from fundamental limitations: low-fidelity kinematics, a lack of dexterous hand gestures, and severely insufficient rich multimodal annotations.
This paper presents Inter-X++, a comprehensive large-scale benchmark captured with a novel hybrid motion capture system. It provides:
- 11,388 high-fidelity interaction sequences with over 8.1 million frames
- Precise whole-body motion and detailed finger articulation
- Rich annotations including:
- Hierarchical fine-grained textual descriptions
- Interaction categories and causal interaction order
- Subject relationships and personalities
- Vertex-level contact maps and physics-based regularization constraints
Leveraging these annotations, the authors formulate a unified testbed covering
four categories of downstream tasks, symmetrically spanning generation and perception paradigms.
They also propose OpenHHI, a unified HHI representation and modeling framework that jointly optimizes interaction reconstruction and semantic understanding. Extensive experiments demonstrate that OpenHHI achieves state-of-the-art performance on both generative and perceptual tasks.
---
*Auto-collected on 2026-08-22*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178633798