论文概要
研究领域: CV
作者: Fidel Omar Tito Cruz, Angie Sanchez Marquina, Summy Farfan, Gissella Bejarano
发布时间: 2026-08-28
arXiv: 2608.28568
中文摘要
手语生成(SLP)旨在从口语生成连续的手势动作,通常通过词汇到姿势的生成实现。先前工作主要遵循两种范式。生成模型从学习到的先验或噪声中合成动作,不参考观察到的手语实例,使得罕见的手部配置和说话者特定的发音难以保留。基于检索的方法重用真实的、发音良好的动作片段,但拼接来自不同说话者和协同发音语境的片段可能在整个序列中引入节奏和风格不一致,不仅在片段边界处。这些局限性表明需要一个互补的解决方案:使用检索提供真实的发音,使用学习到的精修来施加检索单独缺乏的全局连贯性。因此,我们提出了检索-精修范式,从真实检索到的动作开始并将其精修为全局连贯的手语序列,而非从头生成动作。我们的框架SignRR从真实手语片段词典初始化动作,并使用部件感知的残差VQ-VAE精修完整序列,其中残差量化保留精细的手部发音,时间长度差异在潜在空间中处理。在PHOENIX14T和CSL-Daily上的实验表明,SignRR在保持竞争性的姿势质量的同时实现了最先进的回译性能。
原文摘要
Sign language production (SLP) aims to generate continuous signing motion from spoken language, often through gloss-to-pose generation. Prior work mainly follows two paradigms. Generative models synthesize motion from a learned prior or from noise, without reference to an observed signing instance, making rare hand configurations and signer-specific articulation difficult to preserve. Retrieval-based methods reuse real, well-articulated motion segments, but concatenating segments from different signers and co-articulation contexts can introduce rhythm and style inconsistencies across the full sequence, not only at segment boundaries. These limitations suggest a complementary solution: use retrieval to provide realistic articulation, and use learned refinement to impose the global coherence...
自动采集于 2026-09-01
#论文 #arXiv #CV #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。