[论文] SemMSA: Latent Semantic-Aided Robust Multimodal Sentiment Analysis wit...
研究领域: NLP 作者: Wenhao Li, Zhibin Wu, Chong Xiao, Qiangchang Wang 发布时间: 2026-09-24 arXiv: 2609.30238
论文概要
研究领域: NLP 作者: Wenhao Li, Zhibin Wu, Chong Xiao, Qiangchang Wang 发布时间: 2026-09-24 arXiv: 2609.30238
中文摘要
近期关于多模态情感分析(MSA)的研究聚焦于在不完整数据下,从语言、视觉和声学模态推断人类情感。大多数研究通常通过重建模态特征或设计复杂的融合机制来补偿缺失信息。然而,由于部分观测的多模态证据缺乏高层语义基础,这些方法仍然面临虚假生成和噪声引导的问题。为解决这些问题,我们提出 SemMSA——一种潜在语义辅助框架,利用 LLM 构建丰富的情感相关语义,并通过无锚点谱对齐与所有模态充分整合。该方法主要由跨模态语义精炼(CSR)和跨模态谱对齐(CSA)组成。具体而言,CSR 首先通过相应的适配器自适应提取视觉和声学表征,在冻结的 LLM 嵌入空间中与语言构成统一的多模态前缀,然后通过一个 token 高效的潜在精炼过程迭代生成连续的判别性语义状态,无需解码显式文本。接下来,CSA 通过增强各模态核 Gram 矩阵的主导谱分量,将精炼后的语义与所有模态同时对齐。这在不依赖预定义锚点模态的情况下,捕捉了所有表征之间的全局非线性依赖关系。此外,实例级谱分离约束保持了跨样本的判别性并缓解了表征坍缩。在 SIMS、MOSI 和 MOSEI 基准上的大量实验表明,SemMSA 达到了最先进的性能。
原文摘要
Recent research on Multimodal Sentiment Analysis (MSA) has focused on learning from language, visual, and acoustic modalities with incomplete data to infer human sentiment. Most studies typically compensate for missing information by reconstructing modality features or designing complicated fusion mechanisms. However, these methods still suffer from spurious generation and noisy guidance due to the lack of high-level semantic grounding in partially observed multimodal evidence. To address these issues, we propose SemMSA, a latent semantic-aided framework that constructs rich sentiment-relevant semantics with LLMs, fully integrating with all modalities via anchor-free spectral alignment. It mainly consists of Cross-modal Semantic Refinement (CSR) and Cross-modal Spectral Alignment (CSA). Sp...
*自动采集于 2026-09-28*
#论文 #arXiv #NLP #小凯