| 本文已被:浏览 2220次 下载 3225次 |
 码上扫一扫! |
|
|
| 离散语音情感识别研究进展 |
|
郭丽丽1, 王龙标2,3, 党建武2,3,4, 丁世飞1
|
|
1.中国矿业大学 计算机科学与技术学院, 江苏 徐州 221116;2.天津大学 智能与计算学部, 天津 300350;3.天津市认知计算与应用重点实验室 (天津大学), 天津 300350;4.Japan Advanced Institute of Science and Technology, Ishikawa 9231292, Japan
|
|
| 摘要: |
| 语音情感识别是情感计算的重要组成部分, 在人机交互中占据重要的地位. 准确地识别说话人的情感信息, 有助于机器更好地理解用户的意图, 进而提供良好的交互性以提升用户的体验. 以离散语音情感为对象, 对语音情感识别的理论和方法进行综述. 首先在全面回顾情感识别发展历程的同时, 提出一个语音情感识别综述框架. 其次, 介绍情感描述方法以及常用的情感语料库, 旨在为语音情感识别提供基础支撑. 然后, 概述语音情感识别过程, 主要包括特征提取和识别模型, 重点归纳总结传统分类模型、经典深度模型、其他先进模型, 并介绍常用的评价指标, 同时基于评价指标对模型进行总结. 最后, 探讨语音情感识别领域所面临的挑战, 并对未来的发展趋势进行展望. |
| 关键词: 语音情感识别 声学特征 相位信息 分类模型 深度学习 |
| DOI:10.13328/j.cnki.jos.007232 |
| 分类号: |
| 基金项目:国家自然科学基金(62276265, 62176182, 62276185); 中央高校基本科研业务费专项资金(2022QN1096) |
|
| Research Progress of Discrete Speech Emotion Recognition |
|
GUO Li-Li1, WANG Long-Biao2,3, DANG Jian-Wu2,3,4, DING Shi-Fei1
|
|
1.School of Computer Science and Technology, China University of Mining and Technology, Xuzhou 221116, China;2.College of Intelligence and Computing, Tianjin University, Tianjin 300350, China;3.Tianjin Key Laboratory of Cognitive Computing and Application (Tianjin University), Tianjin 300350, China;4.Japan Advanced Institute of Science and Technology, Ishikawa 9231292, Japan
|
| Abstract: |
| Speech emotion recognition is an important part of affective computing and plays an important role in human-computer interaction. Accurately distinguishing emotions helps machines understand users’ intentions and provide better interactivity to enhance user experience. This study reviews the theories and methods of speech emotion recognition focusing on discrete speech emotions. Firstly, the study reviews the development of emotion recognition and presents an architecture of speech emotion recognition to summarize research progress. Secondly, emotion representation models and commonly used corpora are introduced to provide basic support for speech emotion recognition. Then, the process of speech emotion recognition is outlined, including feature extraction and recognition models, with a focus on traditional classification models, classical deep models, and other advanced models. Meanwhile, commonly used evaluation indicators are introduced and applied to provide a summary of models. Finally, the study discusses the challenges in speech emotion recognition and suggests possible directions for future research. |
| Key words: speech emotion recognition (SER) acoustic feature phase information classification model deep learning |