Abstract:Improving robustness to feature noise is one of the most challenging problems in multimodal sentiment analysis (MSA). Recent studies have proposed efficient MSA models that consider missing modalities, but they typically focus on specific types of defects, which leads to limited robustness and generalization in real-world scenarios where multiple types of noise coexist. In addition, deep interactions across different modalities are often inadequately modeled, making it difficult to fully capture multimodal emotional semantics. To address these issues, this study proposes a robust MSA method that integrates adversarial learning with a multi-view network, namely ALMV. Specifically, temporal modality feature masking is adopted to simulate noisy data, and noise-original instance pairs are constructed with intact sequences for data augmentation. Secondly, temporal convolutional networks and Transformer encoders are used to extract local and global information from every modality sequence, and a weight-controlled multi-view network is constructed to learn joint multimodal representations from the noise-original instance pairs. Additionally, a novel multi-level adversarial training strategy with semantic reconstruction supervision is introduced to learn unified representations between noisy and complete data at both the modality level and the utterance level. Extensive experiments on public datasets are conducted to verify the performance of ALMV under various heterogeneous noise scenarios. Experimental results demonstrate that ALMV substantially improves the robustness and performance of multimodal sentiment analysis in the presence of diverse data defects.