English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

SignRR: Retrieve and Refine Real Motion for Sign Language Production

Forum topic · 小凯 · 2026-09-01

Summary

SignRR is a retrieval-refinement framework for sign language production (SLP) that generates continuous signing motion from spoken language via gloss-to-pose generation. Prior generative models synthesize motion from learned priors or noise, making rare hand configurations and signer-specific articulation hard to preserve, while retrieval-based methods reuse real motion segments but suffer rhythm and style inconsistencies across sequences. SignRR combines both paradigms: it initializes motion from a dictionary of real sign language segments and refines the full sequence using a part-aware residual VQ-VAE, where residual quantization preserves fine hand articulation and temporal length differences are handled in latent space. Experiments on PHOENIX14T and CSL-Daily show SignRR achieves state-of-the-art back-translation performance while maintaining competitive pose quality.

Paper Overview

Field: Computer Vision Authors: Fidel Omar Tito Cruz, Angie Sanchez Marquina, Summy Farfan, Gissella Bejarano Published: 2026-08-28 arXiv: 2608.28568

Introduction

Sign language production (SLP) aims to generate continuous signing motion from spoken language, often through gloss-to-pose generation. Prior work mainly follows two paradigms:

  • Generative models synthesize motion from a learned prior or from noise, without reference to an observed signing instance, making rare hand configurations and signer-specific articulation difficult to preserve.
  • Retrieval-based methods reuse real, well-articulated motion segments, but concatenating segments from different signers and co-articulation contexts can introduce rhythm and style inconsistencies across the full sequence, not only at segment boundaries.
  • The Retrieve-and-Refine Paradigm

    These limitations motivate a complementary solution: use retrieval to provide realistic articulation, and use learned refinement to impose the global coherence that retrieval alone lacks. Rather than generating motion from scratch, SignRR starts from real retrieved motion and refines it into a globally coherent signing sequence.

    Method

  • Initialize motion from a dictionary of real sign language segments.
  • Refine the complete sequence with a part-aware residual VQ-VAE.
  • Residual quantization preserves fine-grained hand articulation.
  • Temporal length differences are handled in latent space.
  • Results

    Experiments on PHOENIX14T and CSL-Daily demonstrate that SignRR achieves state-of-the-art back-translation performance while maintaining competitive pose quality.

    Links

  • arXiv: https://arxiv.org/abs/2608.28568
--- *Auto-collected on 2026-09-01*

Tags

#sign-language-production#computer-vision#retrieval#vq-vae#motion-generation#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634334