| 本文已被:浏览 3842次 下载 3424次 |
 码上扫一扫! |
|
|
| 面向深度学习的图像数据增强综述 |
|
杨锁荣1,2, 杨洪朝1,2, 申富饶1,3, 赵健4
|
|
1.计算机软件新技术国家重点实验室(南京大学), 江苏 南京 210023;2.南京大学 计算机科学与技术系, 江苏 南京 210023;3.南京大学 人工智能学院, 江苏 南京 210023;4.南京大学 电子科学与工程学院, 江苏 南京 210023
|
|
| 摘要: |
| 深度学习已经在许多计算机视觉任务中取得了显著的成果. 然而, 深度神经网络通常需要大量的训练数据以避免过拟合, 但实际应用中标记数据可能非常有限. 因此, 数据增强已成为提高训练数据充分性和多样性的有效方法, 也是深度学习模型成功应用于图像数据的必要环节. 系统地回顾不同的图像数据增强方法, 并提出一个新的分类方法, 为研究图像数据增强提供了新的视角. 从不同的类别出发介绍各类数据增强方法的优势和局限性, 并阐述各类方法的解决思路和应用价值. 此外, 还介绍语义分割、图像分类和目标检测这3种典型计算机视觉任务中常用的公共数据集和性能评价指标, 并在这3个任务上对数据增强方法进行实验对比分析. 最后, 讨论当前数据增强所面临的挑战和未来的发展趋势. |
| 关键词: 深度学习 图像数据增强 图像识别 泛化性能 计算机视觉 |
| DOI:10.13328/j.cnki.jos.007263 |
| 分类号: |
| 基金项目:国家自然科学基金(62276127) |
|
| Image Data Augmentation for Deep Learning: A Survey |
|
YANG Suo-Rong1,2, YANG Hong-Chao1,2, SHEN Fu-Rao1,3, ZHAO Jian4
|
|
1.State Key Laboratory for Novel Software Technology (Nanjing University), Nanjing 210023, China;2.Department of Computer Science and Technology, Nanjing University, Nanjing 210023, China;3.School of Artificial Intelligence, Nanjing University, Nanjing 210023, China;4.School of Electronic Science and Engineering, Nanjing University, Nanjing 210023, China
|
| Abstract: |
| Deep learning has yielded remarkable achievements in many computer vision tasks. However, deep neural networks typically require a large amount of training data to prevent overfitting. In practical applications, labeled data may be extremely limited. Thus, data augmentation has become an effective way to enhance the adequacy and diversity of training data and is also a necessary link for the successful application of deep learning models to image data. This study systematically reviews different image data augmentation methods and proposes a new classification method to provide a fresh perspective for studying image data augmentation. The advantages and limitations of various data augmentation methods are introduced from different categories, and the solution ideas and application values of these methods are elaborated. In addition, commonly used public datasets and performance evaluation indicators in three typical computer vision tasks of semantic segmentation, image classification, and object detection are presented. Experimental comparative analysis of data augmentation methods is conducted on these three tasks. Finally, the challenges and future development trends currently faced by data augmentation are discussed. |
| Key words: deep learning image data augmentation image recognition generalization performance computer vision |