基于非纯测量场景的隐变量因果结构学习算法
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

TP18

基金项目:

NSFC-区域创新发展联合基金(广东) (U24A20233)


Causal Structure Learning Algorithm for Latent Variables in Impure Measurement Scenarios
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    隐变量间的因果结构学习, 核心在于从观测数据中挖掘隐变量彼此的因果关联, 是科学研究中揭示现象本质的一种关键技术. 现有方法普遍依赖“纯子”假设, 即隐变量对应的观测子代变量(测量变量)间不存在直接的因果连接边, 该假设在许多现实场景中往往无法成立, 从而导致现有方法可识别性受限. 针对该挑战, 考虑“非纯”测量场景下的隐变量识别性问题, 提出线性非高斯无环隐变量模型(linear non-Gaussian acyclic latent variable model, LiNGLM), 该模型允许观测变量间存在因果关联. 基于该模型, 提出一种隐变量结构学习(latent variable structure learning, LLSTIN)算法, 该算法基于变换独立噪声(transformed independent noise, TIN)条件及其建立的图准则, 首先通过因果聚类构建(build causal cluster, BCC)算法提取各隐变量对应的“单因子集”来识别隐变量的存在, 然后选取集合中的“根”观测变量作为隐变量的有效代理, 进而识别隐变量间的因果关系. 从理论上证明算法可正确提取单因子集并进一步识别隐变量因果结构, 仿真数据与真实数据实验结果进一步验证所提算法的正确性和有效性.

    Abstract:

    Learning the causal structure among latent variables is a key technique for revealing the underlying mechanisms of phenomena in scientific research, with its core objective being to infer causal relationships between latent variables from observational data. Existing methods generally rely on the “pure child” assumption, which assumes that no direct causal connections exist among the observed child variables (measurement variables) corresponding to latent variables. However, this assumption often fails to hold in many real-world scenarios, thus limiting the identifiability of existing methods. To address this challenge, this study investigates the identifiability of latent variables under the “non-pure” measurement setting and proposes a linear non-Gaussian acyclic latent variable model (LiNGLM) that allows causal relationships among observed variables. Based on this model, a latent variable structure learning (LLSTIN) algorithm is proposed. This algorithm is based on the transformed independent noise (TIN) condition and the graph criterion established thereby. First, the “single-factor sets” corresponding to each latent variable are extracted through the build causal cluster (BCC) algorithm to identify the existence of latent variables. Then, the “root” observed variables in the sets are selected as effective proxies for latent variables, and the causal relationships between latent variables are further identified. This study theoretically proves that the proposed algorithm can correctly extract single-factor sets and further identify the causal structure among latent variables. Simulation data and real-world data experimental results further verify the correctness and effectiveness of the proposed algorithm.

    参考文献
    相似文献
    引证文献
引用本文

王馥弘,陈正鸣,夏业伟,乔杰,郝志峰,蔡瑞初.基于非纯测量场景的隐变量因果结构学习算法.软件学报,,():1-19

复制
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-10-14
  • 最后修改日期:2025-11-11
  • 录用日期:
  • 在线发布日期: 2026-05-20
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62562563 传真:010-62562533 Email:jos@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号