引用本文:孙伟松,陈宇琛,赵梓含,陈宏,葛一飞,韩廷旭,黄胜寒,李佳讯,房春荣,陈振宇.深度代码模型安全综述.软件学报,2025,36(4):1461-1488
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览 2991次   下载 2664 本文二维码信息
码上扫一扫!
分享到: 微信 更多
深度代码模型安全综述
孙伟松1,2, 陈宇琛1,2, 赵梓含1,2, 陈宏1,2, 葛一飞1,2, 韩廷旭1,2, 黄胜寒1,2, 李佳讯3, 房春荣1,2, 陈振宇1,2
1.计算机软件新技术国家重点实验室 (南京大学), 江苏 南京 210093;2.南京大学 软件学院, 江苏 南京 210093;3.苏州大学 数学科学学院, 江苏 苏州 215006
摘要:
随着深度学习技术在计算机视觉与自然语言处理等领域取得巨大成功, 软件工程研究者开始尝试将其引入到软件工程任务求解当中. 已有研究结果显示, 深度学习技术在各种代码相关任务(例如代码检索与代码摘要)上具有传统方法与机器学习方法无法比拟的优势. 这些面向代码相关任务训练的深度学习模型统称为深度代码模型. 然而, 由于神经网络的脆弱性和不可解释性, 与自然语言处理模型与图像处理模型一样, 深度代码模型安全也面临众多挑战, 已经成为软件工程领域的焦点. 近年来, 研究者提出了众多针对深度代码模型的攻击与防御方法. 然而, 目前仍缺乏对深度代码模型安全研究的系统性综述, 不利于后续研究者对该领域进行快速的了解. 因此, 为了总结该领域研究现状、挑战及时跟进该领域的最新研究成果, 搜集32篇该领域相关论文, 并将现有的研究成果主要分为后门攻击与防御技术和对抗攻击与防御技术两类. 按照不同技术类别对所收集的论文进行系统地梳理和总结. 随后, 总结该领域中常用的实验数据集和评估指标. 最后, 分析该领域所面临的关键挑战以及未来可行的研究方向, 旨在为后续研究者进一步推动深度代码模型安全的发展提供有益指导.
关键词:  深度代码模型  深度代码模型安全  人工智能模型安全  后门攻击与防御  对抗攻击与防御
DOI:10.13328/j.cnki.jos.007254
分类号:
基金项目:国家自然科学基金(61932012, 62372228)
Survey on Security of Deep Code Models
SUN Wei-Song1,2, CHEN Yu-Chen1,2, ZHAO Zi-Han1,2, CHEN Hong1,2, GE Yi-Fei1,2, HAN Ting-Xu1,2, HUANG Sheng-Han1,2, LI Jia-Xun3, FANG Chun-Rong1,2, CHEN Zhen-Yu1,2
1.State Key Laboratory for Novel Software Technology (Nanjing University), Nanjing 210093, China;2.Software Institute, Nanjing University, Nanjing 210093, China;3.School of Mathematical Sciences, Soochow University, Suzhou 215006, China
Abstract:
With the significant success of deep learning in fields such as computer vision and natural language processing, researchers in software engineering have begun to explore its integration into solving software engineering tasks. Existing research indicates that deep learning exhibits advantages in various code-related tasks, such as code retrieval and code summarization, that traditional methods and machine learning cannot match. Deep learning models trained for code-related tasks are referred to as deep code models. However, similar to natural language processing and image processing models, the security of deep code models faces numerous challenges due to the vulnerability and inexplicability of neural networks. It has become a research focus in software engineering. In recent years, researchers have proposed numerous attack and defense methods for deep code models. Nevertheless, there is a lack of a systematic review of research on deep code model security, hindering the rapid understanding of subsequent researchers in this field. To provide a comprehensive overview of the current research, challenges, and latest findings in this field, this study collects 32 relevant papers and categorizes existing research results into two main classes: backdoor attack and defense techniques, and adversarial attack and defense techniques. This study systematically analyzes and summarizes the collected papers based on the above two categories. Subsequently, it outlines commonly used experimental datasets and evaluation metrics in this field. Finally, it analyzes key challenges in this field and suggests feasible future research directions, aiming to provide valuable guidance for further advancements in the security of deep code models.
Key words:  deep code model  security of deep code model  security of artificial intelligence model  backdoor attack and defense  adversarial attack and defense

引用本文:
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览次   下载  
分享到: 微信 更多
摘要:
关键词:  
DOI:
分类号:
基金项目:
Abstract:
Key words: