English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Inter-X++: A Comprehensive Benchmark for Multimodal Human-Human Interaction

Forum topic · 小凯 · 2026-08-22

Summary

Inter-X++ (arXiv:2608.20312) is a large-scale benchmark for multimodal human-human interaction (HHI) research in computer vision. Captured with a novel hybrid motion capture system, it provides 11,388 high-fidelity interaction sequences totaling over 8.1 million frames, featuring precise full-body motion and detailed finger articulation. The dataset addresses fundamental limitations of prior datasets, which suffered from low-fidelity kinematics, missing dexterous hand gestures, and insufficient multimodal annotation. Each sequence is enriched with hierarchical fine-grained textual descriptions, interaction categories, causal interaction ordering, subject relationships and personalities, vertex-level contact maps, and physics-regularized constraints. Building on these annotations, the authors establish a unified testbed spanning four downstream task families that symmetrically cover both generation and perception paradigms. They also introduce OpenHHI, a unified HHI representation and modeling framework that jointly optimizes interaction reconstruction and semantic understanding, achieving state-of-the-art performance on both generation and perception benchmarks.

Paper Overview

  • Field: Computer Vision (CV)
  • Authors: Liang Xu, Chengqun Yang, Zili Lin, Xintao Lv
  • Published: 2026-08-22
  • arXiv: 2608.20312
  • Abstract

    The ability to perceive and synthesize human-human interaction is fundamental to developing intelligent digital human systems. However, existing datasets and modeling approaches are severely limited by low-fidelity kinematics, a lack of dexterous hand gestures, and insufficient rich multimodal annotation.

    This paper presents Inter-X++, a comprehensive large-scale benchmark captured with a novel hybrid motion capture system. It provides 11,388 high-fidelity interaction sequences and over 8.1 million frames, featuring precise full-body motion and detailed finger articulation.

    The dataset is enriched with multi-faceted annotations, including:

  • Hierarchical fine-grained textual descriptions
  • Interaction categories
  • Causal interaction ordering
  • Subject relationships and personalities
  • Vertex-level contact maps
  • Physics-regularized constraints

Contributions

1. Unified testbed: Using these annotations, the authors define four families of downstream tasks, symmetrically spanning both generation and perception paradigms. 2. OpenHHI framework: A unified HHI representation and modeling framework that jointly optimizes interaction reconstruction and semantic understanding. 3. State-of-the-art results: Extensive experiments show that OpenHHI achieves state-of-the-art performance across both generation and perception tasks.

---

*Source link: https://arxiv.org/abs/2608.20312*

Tags

#computer-vision#human-human-interaction#benchmark#motion-capture#dataset#multimodal#motion-generation#paper

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633819