Journal of Software:2020.31(6):1747-1760

(电子科技大学 计算机科学与工程学院, 四川 成都 611731)
Chinese Sentence-Level Lip Reading Based on End-to-End Model
ZHANG Xiao-Bing,GONG Hai-Gang,YANG Fan,DAI Xi-Li
(School of Computer Science and Engineering, University of Electronic Science and Technology of China, Chengdu 611731, China)
Chart / table
Similar Articles
Article :Browse 143   Download 169
Received:May 10, 2018    Revised:September 04, 2018
> 中文摘要: 近年来,随着深度学习的广泛应用,唇语识别技术也取得了快速的发展.与传统的方法不同,在基于深度学习的唇语识别模型中,通常包含使用神经网络对图像进行特征提取和特征理解两个部分.根据中文唇语识别的特点,将识别过程划分为两个阶段——图片到拼音(P2P)以及拼音到汉字(P2CC)的识别.分别设计两个不同子网络针对不同的识别过程,当两个子网络训练好后,再把它们放在一起进行端到端的整体架构优化.由于目前没有可用的中文唇语数据集,因此采用半自动化的方法从CCTV官网上收集了6个月20.95GB的中文唇语数据集CCTVDS,共包含14 975个样本.此外,额外采集了269 558条拼音汉字样本数据对拼音到汉字识别模块进行预训练.在CCTVDS数据集上的实验结果表明,所提出的ChLipNet可分别达到45.7%的句子识别准确率和58.5%的拼音序列识别准确率.此外,ChLipNet不仅可以加速训练、减少过拟合,并且能够克服汉语识别中的歧义模糊性.
Abstract:In recent years, with the widely application of deep learning, lip reading recognition technology has achieved rapid development. Different from traditional methods, lip reading recognition methods based on the deep learning usually use the neural network model both for the feature extraction and comprehension. According to the characteristics of Chinese language, a two-step end-to-end architecture is implemented, in which two deep neural network modules are applied to perform the recognition of picture-to-pinyin (P2P) and pinyin-to-hanzi (P2CC) respectively. After the two modules are trained with convergence, they are then jointly optimized to improve the overall performance. Due to the lack of Chinese lip reading dataset, the 6-month daily news broadcasts are collected from China Central Television (CCTV), and they are semi-automatically labelled into a 20.95 GB dataset CCTVDS with 14 975 samples. In addition, the supplementary dataset with 269 558 samples are collected during the pre-training of P2CC. According to experimental results trained on the CCTVDS, the proposed ChLipNet can achieve 45.7% sentence-level and 58.5% Pinyin-level accuracies. In addition, ChLipNet can not only accelerate training, reduce overfitting, but also overcome syntactic ambiguity in the recognition of Chinese language.
文章编号:     中图分类号:TP18    文献标志码:
基金项目:国家自然科学基金(61572113) 国家自然科学基金(61572113)
Foundation items:National Natural Science Foundation of China (61572113)
Reference text:


ZHANG Xiao-Bing,GONG Hai-Gang,YANG Fan,DAI Xi-Li.Chinese Sentence-Level Lip Reading Based on End-to-End Model.Journal of Software,2020,31(6):1747-1760