融合对抗学习与多视图网络的鲁棒多模态情感分析
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

国家自然科学基金(62277008); 重庆市自然科学基金创新发展联合基金(重点)项目(CSTB2024NSCQ-LZX0133)


Robust Multimodal Sentiment Analysis Integrating Adversarial Learning and Multi-view Network
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    提高模型对特征噪声的鲁棒性已经成为多模态情感分析(multimodal sentiment analysis, MSA)领域最具挑战性的任务之一, 最近的研究已经提出了考虑缺失模态的高效MSA模型. 然而, 现有方法通常只关注特定类型的缺陷, 导致在同时存在多种类型噪声的实际场景中模型鲁棒性和泛化性弱, 而且现有研究忽略了不同模态之间的深层交互, 难以全面理解多模态情感信息. 针对这些问题, 提出一种融合对抗学习与多视图网络的鲁棒多模态情感分析网络(robust multimodal sentiment analysis integrating adversarial learning and multi-view network, ALMV). 具体而言, 首先使用时间模态特征缺失作为噪声数据, 与完整序列构成噪声-原始实例对以构建增强数据. 其次, 通过时间卷积和Transformer编码器提取噪声-原始实例对中每个模态序列的局部信息和全局信息, 并构建权重调控的多视图网络来学习噪声-原始实例对之间的多模态联合表示. 然后, 提出一种具有语义重建监督的多重对抗训练策略, 在模态级和话语级两个层面学习噪声和完整数据之间的统一联合表征. 最后, 在公开数据集上进行了大量实验, 对多种类型噪声场景下ALMV的性能进行测试. 实验结果证明, ALMV方法有效地提升了多模态情感分析在多种异构数据缺陷场景下的性能和鲁棒性.

    Abstract:

    Improving robustness to feature noise is one of the most challenging problems in multimodal sentiment analysis (MSA). Recent studies have proposed efficient MSA models that consider missing modalities, but they typically focus on specific types of defects, which leads to limited robustness and generalization in real-world scenarios where multiple types of noise coexist. In addition, deep interactions across different modalities are often inadequately modeled, making it difficult to fully capture multimodal emotional semantics. To address these issues, this study proposes a robust MSA method that integrates adversarial learning with a multi-view network, namely ALMV. Specifically, temporal modality feature masking is adopted to simulate noisy data, and noise-original instance pairs are constructed with intact sequences for data augmentation. Secondly, temporal convolutional networks and Transformer encoders are used to extract local and global information from every modality sequence, and a weight-controlled multi-view network is constructed to learn joint multimodal representations from the noise-original instance pairs. Additionally, a novel multi-level adversarial training strategy with semantic reconstruction supervision is introduced to learn unified representations between noisy and complete data at both the modality level and the utterance level. Extensive experiments on public datasets are conducted to verify the performance of ALMV under various heterogeneous noise scenarios. Experimental results demonstrate that ALMV substantially improves the robustness and performance of multimodal sentiment analysis in the presence of diverse data defects.

    参考文献
    相似文献
    引证文献
引用本文

蔡林沁,刘道宏,刘岚睿,牟仁云.融合对抗学习与多视图网络的鲁棒多模态情感分析.软件学报,,():1-22

复制
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-07-07
  • 最后修改日期:2025-10-19
  • 录用日期:
  • 在线发布日期: 2026-07-08
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62562563 传真:010-62562533 Email:jos@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号