English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Inter-X++: A Large-Scale Multimodal Benchmark for Human-Human Interaction

Forum topic · 小凯 · 2026-08-22

Summary

Inter-X++ is a comprehensive large-scale benchmark for perceiving and synthesizing human-human interaction (HHI), collected via a novel hybrid motion capture system. It provides 11,388 high-fidelity interaction sequences with over 8.1 million frames, featuring accurate whole-body motion and detailed finger articulation. The dataset is enriched with multifaceted annotations, including hierarchical fine-grained textual descriptions, interaction categories, causal interaction order, subject relationships and personalities, vertex-level contact maps, and physics-based regularization constraints. The authors establish a unified testbed covering four categories of downstream tasks spanning both generation and perception paradigms, and introduce OpenHHI, a unified HHI representation and modeling framework that jointly optimizes interaction reconstruction and semantic understanding. Extensive experiments show OpenHHI achieves state-of-the-art performance across both generative and perceptual tasks. Paper: arXiv 2608.20312.

Paper Overview

Field: Computer Vision (CV) Authors: Liang Xu, Chengqun Yang, Zili Lin, Xintao Lv Published: 2026-08-22 arXiv: 2608.20312

Abstract

The ability to perceive and synthesize human-human interaction (HHI) is fundamental to developing intelligent digital human systems. However, existing datasets and modeling methods suffer from fundamental limitations: low-fidelity kinematics, a lack of dexterous hand gestures, and severely insufficient rich multimodal annotations.

This paper presents Inter-X++, a comprehensive large-scale benchmark captured with a novel hybrid motion capture system. It provides:

  • 11,388 high-fidelity interaction sequences with over 8.1 million frames
  • Precise whole-body motion and detailed finger articulation
  • Rich annotations including:
  • Hierarchical fine-grained textual descriptions
  • Interaction categories and causal interaction order
  • Subject relationships and personalities
  • Vertex-level contact maps and physics-based regularization constraints
Leveraging these annotations, the authors formulate a unified testbed covering four categories of downstream tasks, symmetrically spanning generation and perception paradigms.

They also propose OpenHHI, a unified HHI representation and modeling framework that jointly optimizes interaction reconstruction and semantic understanding. Extensive experiments demonstrate that OpenHHI achieves state-of-the-art performance on both generative and perceptual tasks.

--- *Auto-collected on 2026-08-22*

Tags

#computer-vision#human-human-interaction#benchmark#motion-capture#motion-generation#multimodal#openhhi#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633798