English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Action Motifs: Self-Supervised Hierarchical Representation of Human Behavior (A4Mer)

Forum topic · 小凯 · 2026-05-02

Summary

This paper introduces a hierarchical representation for human behavior modeling built on the compositionality of body movement. Action Atoms capture atomic joint movements, while Action Motifs emerge from their temporal compositions, encoding similar body movements shared across different overall human actions. The authors present A4Mer, a nested latent-space Transformer that learns this hierarchy fully self-supervised from 3D human pose data. A4Mer segments pose sequences into variable-length parts, each represented as a single latent token (Action Atom), and trains via masked token prediction in separate latent spaces; reusable, semantic temporal patterns (Action Motifs) emerge bottom-up. The paper also releases the Action Motif Dataset (AMD), a large-scale multi-view video dataset with complete SMPL annotations, using foot-mounted cameras to enable per-frame labeling despite frequent severe body occlusion. Experiments show A4Mer extracts meaningful Action Motifs that improve action recognition, motion prediction, and motion interpolation.

Paper Overview

Field: Computer Vision (CV) Authors: Zhenyu Wu, Ziyun Wang, Yixuan Wei et al. Published: 2026-04-30 arXiv: 2604.28173

Original Abstract

Effective human behavior modeling requires a representation of the human body movement that capitalizes on its compositionality. We propose a hierarchical representation consisting of Action Atoms that capture the atomic joint movements and Action Motifs that are formed by their temporal compositions and encode similar body movements found across different overall human actions. We derive A4Mer, a nested latent Transformer to learn this hierarchical representation from human pose data in a fully self-supervised manner.

Key Contributions (from the Chinese summary)

  • Hierarchical representation: Action Atoms capture atomic joint movements; Action Motifs are their temporal compositions, encoding reusable, semantic body movement segments shared across different overall actions.
  • A4Mer architecture: a nested latent-space Transformer that segments 3D pose sequences into variable-length segments, each encoded as a single latent token (Action Atom).
  • Self-supervised training: masked token prediction in the respective latent spaces unifies the pretraining tasks; Action Motifs emerge naturally bottom-up.
  • Action Motif Dataset (AMD): a large-scale multi-view human behavior video dataset with complete SMPL annotations. Cameras are mounted on the feet to enable per-frame annotation despite frequent and severe body occlusion.
  • Results: A4Mer extracts meaningful Action Motifs that bring significant improvements to human behavior modeling tasks, including action recognition, motion prediction, and motion interpolation.
---

*Auto-collected on 2026-05-02*

Tags

#computer-vision#self-supervised-learning#human-pose#action-recognition#motion-prediction#transformer#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619039