Int-AD: 面向提示注入攻击的攻防一体化框架
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

国家自然科学基金重点项目(62332004); 四川省揭榜挂帅项目(2024YFCY0003); 成都市产业链协同创新项目(2025-XT00-00017-GX); 中央高校项目(ZYGX2025TS010); 智能系统安全实验室 (Q25008)


Int-AD: Integrated Attack and Defense Framework for Prompt Injection Attack
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    提示注入攻击(prompt injection attack)已成为大语言模型(LLM)应用日益严峻的安全威胁, 位于OWASP开放式Web应用程序安全项目10大威胁之首. 此类攻击通过在提示词中嵌入恶意内容来诱导大模型生成误导性输出, 带来信息泄露、系统滥用等严重安全隐患. 然而, 目前针对此类攻击的攻击策略评估指标仅有攻击成功率, 难以全面评估攻击效果; 在防御策略方面, 现有方法通常只能单独用于预防或检测, 两者难以兼得. 深入分析提示注入攻击和防御的特征及其对大模型输出的影响, 提出攻防一体化框架Int-AD, 包括: 采用情感增强和控制输出的提示注入攻击方法CoA (control-output attack)和同时具备预防和检测能力的防御策略UnD (universal defense). 在攻击方面, 定义两个新的攻击评价指标: 攻击干扰率(attack interference rate, AIR)与攻击误导率(attack misdirection rate, AMR), 用于在攻击结果不同时分别进行精准评估, 完善攻击评估指标. 通过情感增强和控制输出设计提示词, 将AIR和AMR与现有METEOR评估指标联合使用来评估CoA的攻击效率; 在防御方面, UnD兼具预防和检测能力, 可提高防御成功率并检测提示注入攻击的影响. 在Qwen、Llama、DeepSeek这3个主流大模型上采用squad和web_questions两个公开数据集进行实验, CoA攻击方法的平均METEOR得分为0.050, 相比当前最好方法Combine可降低42.5%; 平均AIR和AMR分别达到0.97和0.45, 相比Combine分别提升6.59%和28.57%, 表明CoA具有更强的攻击能力. UnD的平均防御成功率达80.6%, 相比当前最好方法Sandwich提高9.07%, 其平均KMR (knownanswer matching rate)为0.921, 较方法Knownanswer提升58.52%, 表明UnD具有更强的防御能力.

    Abstract:

    Prompt injection attacks have become an increasingly serious security threat to large language models (LLMs) and are ranked first in the OWASP Top 10 for LLM applications. Such attacks induce LLMs to generate misleading outputs by embedding malicious content in prompts, thus posing serious security risks such as information leakage and system abuse. However, current evaluation metrics for attack strategies against such attacks only include the attack success rate, making it difficult to comprehensively assess attack effectiveness. In terms of defense strategies, existing methods can typically be used for either prevention or detection, but it is difficult to achieve both simultaneously. This study presents an in-depth analysis of the characteristics of prompt injection attacks and defenses, as well as their impact on LLM outputs, and proposes an integrated attack-defense framework termed Int-AD. The proposed framework includes a prompt injection attack method, control-output attack (CoA), which employs sentiment enhancement and output control, and a defense strategy, universal defense (UnD), which possesses both prevention and detection capabilities. In terms of attacks, two new evaluation metrics are defined: attack interference rate (AIR) and attack misdirection rate (AMR), which are used to conduct precise evaluations under different attack outcomes, thus improving the evaluation metrics. By designing prompts with sentiment enhancement and output control, AIR and AMR are combined with the existing METEOR evaluation metric to assess the attack efficiency of CoA. On the defense side, UnD possesses both prevention and detection capabilities, which improves the defense success rate and enables the detection of the impact of prompt injection attacks. Experiments are conducted on three mainstream large models, Qwen, Llama, and DeepSeek, using the squad and web_questions public datasets. The average METEOR score of the CoA attack method is 0.050, which is 42.5% lower than that of the current best method, Combine. The average AIR and AMR reach 0.97 and 0.45, respectively, representing improvements of 6.59% and 28.57% over Combine and indicating that CoA possesses stronger attack capability. The average defense success rate of UnD is 80.6%, which is 9.07% higher than that of the current best method, Sandwich. Its average knownanswer matching rate (KMR) is 0.921, which is 58.52% higher than that of the Knownanswer method, indicating that UnD possesses stronger defense capability.

    参考文献
    相似文献
    引证文献
引用本文

程翔,曹晟,陈厅,李雄,张小松. Int-AD: 面向提示注入攻击的攻防一体化框架.软件学报,,():1-17

复制
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-09-03
  • 最后修改日期:2026-03-06
  • 录用日期:
  • 在线发布日期: 2026-07-01
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62562563 传真:010-62562533 Email:jos@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号