大语言模型中逻辑神经元的系统性探测与影响分析
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

TP18

基金项目:

国家自然科学基金 (62376270); 北京市新一代信息通信技术创新专项(Z251100008125025)


Systematic Probing and Impact Analysis of Logic Neurons in Large Language Model
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    大语言模型(large language model, LLM)在涉及逻辑操作的复杂推理任务中表现良好, 然而其逻辑推理能力的内部机理仍缺乏明确的可解释性分析. 目前尚不明确LLM内部是否存在与逻辑推理任务表现密切相关的神经元子集, 这些神经元在网络中以何种形式进行计算表征, 它们又如何影响模型在不同任务中的表现. 为探究上述涉及LLM推理机制本质的关键问题, 提出一种“识别-干预-评估”框架, 旨在系统性地剖析逻辑推理能力的神经元运行基础与功能实现. 利用基于梯度重要性的神经元定位方法, 定位对逻辑推理任务高度敏感的“逻辑神经元”, 并结合激活值进行定向干预, 进而评估其对模型行为的影响. 研究结果表明, 所定位的逻辑相关神经元子集主要集中于网络中层, 并以前馈网络模块(feed-forward network, FFN)为主导. 进一步的分析表明, 不同逻辑任务之间既存在共享神经元成分, 也存在任务特异性差异. 基于这些发现, 探索一种无需修改模型权重、轻量级的逻辑推理能力定向增强方法, 实验结果显示, 该方法在LogicBench的长链推理任务上将模型平均准确率提升9.2%, 并在逻辑敏感型通用任务上性能提升4.57%.

    Abstract:

    Large language models (LLMs) perform well in complex reasoning tasks involving logical operations. However, the internal mechanisms underlying their logical reasoning capabilities remain insufficiently interpretable. Currently, it remains unclear whether neuron subsets closely associated with logical reasoning performance exist within LLMs, what forms of computational operations are performed by these neurons in the network, and how they influence model performance across different tasks. To investigate these key questions concerning the nature of LLM reasoning mechanisms, this study proposes an “identification-intervention-evaluation” framework aimed at systematically analyzing the neuronal operational basis and functional implementation of logical reasoning. A neuron localization method based on gradient importance is used to identify “logic neurons” that are highly sensitive to logical reasoning tasks, and targeted interventions are performed on their activation values to further evaluate their impact on model behavior. The results show that the identified logic-related neuron subsets are mainly concentrated in the middle layers of the network and predominantly located in feed-forward network (FFN) modules. Further analysis indicates that different logical tasks involve both shared neuron components and task-specific differences. Based on these findings, this study explores a lightweight targeted enhancement method for logical reasoning capabilities without modifying model weights. Experimental results show that this method improves the average accuracy of the model on long-chain reasoning tasks in LogicBench by 9.2% and improves performance on logic-sensitive general tasks by 4.57%.

    参考文献
    相似文献
    引证文献
引用本文

姜慧,何世柱,刘康,赵军.大语言模型中逻辑神经元的系统性探测与影响分析.软件学报,,():1-22

复制
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-02-10
  • 最后修改日期:2026-04-29
  • 录用日期:
  • 在线发布日期: 2026-08-12
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62562563 传真:010-62562533 Email:jos@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号