| 引用本文: | 王维靖,陈俊洁,杨林,侯德俊,王星凯,吴复迪,张润滋,王赞.基于多元数据融合的网络侧告警排序方法.软件学报,2024,35(8):3610-3625 |
| |
|
| |
|
|
| 本文已被:浏览 1737次 下载 4057次 |
 码上扫一扫! |
|
|
| 基于多元数据融合的网络侧告警排序方法 |
|
王维靖1,2, 陈俊洁1, 杨林1, 侯德俊3, 王星凯4, 吴复迪4, 张润滋4, 王赞1,2
|
|
1.天津大学智能与计算学部, 天津 300350;2.天津大学新媒体与传播学院, 天津 300072;3.天津大学信息与网络中心, 天津 300072;4.绿盟科技集团股份有限公司, 北京 100089
|
|
| 摘要: |
| 部署在网络节点上的网络安全监控系统每天会生成海量网络侧告警, 导致安全人员面临巨大压力, 并使其对高风险告警不再敏感, 无法及时发现网络攻击行为. 由于网络攻击行为的复杂多变以及网络侧告警信息的局限性, 已有面向IT运维的告警排序/分类方法并不适用于网络侧告警. 因此, 提出了基于多元数据融合的首个网络侧告警排序方法NAP (network-side alert prioritization). NAP首先设计了一个基于源IP地址与目的IP地址的多策略上下文编码器, 用于捕获告警的上下文信息; 其次, NAP设计了一个基于注意力机制双向GRU (gated recurrent unit)模型与ChineseBERT模型的文本编码器, 从告警报文等文本数据中学习网络侧告警的语义信息; 最后, NAP构建了排序模型得到告警排序值, 并按其降序将攻击性强的高风险告警排在前面, 从而优化网络侧告警管理流程. 在3组绿盟科技网络攻防数据上的实验表明: NAP能够有效且稳定地排序网络侧告警, 并且显著优于对比方法. 例如: 平均排序指标NDCG@k (kÎ[1,10]) (即前1-10个排序结果的归一化折损累计增益)均在0.893 1-0.958 3之间, 比最先进的方法提升64.73%以上. 另外, 通过将NAP应用于天津大学真实的网络侧告警数据, 进一步证实了其实用性. |
| 关键词: 网络安全 网络侧告警 排序 数据融合 |
| DOI:10.13328/j.cnki.jos.007118 |
| 分类号: |
| 基金项目:北京市科技新星计划(Z211100002121150) |
|
| Network-side Alert Prioritization Method Based on Multivariate Data Fusion |
|
WANG Wei-Jing1,2, CHEN Jun-Jie1, YANG Lin1, HOU De-Jun3, WANG Xing-Kai4, WU Fu-Di4, ZHANG Run-Zi4, WANG Zan1,2
|
|
1.College of Intelligence and Computing, Tianjin University, Tianjin 300350, China;2.School of New Media and Communication, Tianjin University, Tianjin 300072, China;3.Information and Network Center, Tianjin University, Tianjin 300072, China;4.NSFOCUS Technologies Group Co. Ltd., Beijing 100089, China
|
| Abstract: |
| The network security monitoring systems deployed on network nodes generate a large number of network-side alerts every day, causing the security engineers to face significant pressure to lose sensitivity to high-risk alerts and fail to detect network attacks in time. Due to the complexity and variability of cyber attacks and the limitation of network-side alert information, existing alert prioritization/ classification methods for IT operations are unsuitable for network-side alerts. Thus, network-side alert prioritization (NAP), the first network-side alert prioritization method, is proposed based on multivariate data fusion. NAP first designs a multi-strategy context encoder based on source IP address and destination IP address to capture the context information of network-side alerts. And then, NAP designs a text encoder based on the attention-based bidirectional GRU model and the ChineseBERT model to learn the semantic information of network-side alerts from the text data such as alert messages. Finally, NAP builds a ranking model to obtain the alert ranking values and then ranks the high-risk alerts with cyber attack intention in the front according to their descending order to optimize the network-side alert management process. The experiments on three groups of network attack and defense data from NSFOCUS show that NAP can achieve effective and stable prioritization results, and significantly outperforms the compared methods. For example, the average NDCG@k (kÎ[1,10]) (i.e., normalized discounted cumulative gain of the first 1 to 10 ranking results) ranges from 0.893 1 to 0.958 3, and outperforms the state-of-the-art method more than 64.73%. Besides, NAP has been applied to a real-world network-side alert dataset from Tianjin University, further confirming its practicability. |
| Key words: cyber security network-side alert prioritization data fusion |
|
|
|
|