Abstract:Learning the causal structure among latent variables is a key technique for revealing the underlying mechanisms of phenomena in scientific research, with its core objective being to infer causal relationships between latent variables from observational data. Existing methods generally rely on the “pure child” assumption, which assumes that no direct causal connections exist among the observed child variables (measurement variables) corresponding to latent variables. However, this assumption often fails to hold in many real-world scenarios, thus limiting the identifiability of existing methods. To address this challenge, this study investigates the identifiability of latent variables under the “non-pure” measurement setting and proposes a linear non-Gaussian acyclic latent variable model (LiNGLM) that allows causal relationships among observed variables. Based on this model, a latent variable structure learning (LLSTIN) algorithm is proposed. This algorithm is based on the transformed independent noise (TIN) condition and the graph criterion established thereby. First, the “single-factor sets” corresponding to each latent variable are extracted through the build causal cluster (BCC) algorithm to identify the existence of latent variables. Then, the “root” observed variables in the sets are selected as effective proxies for latent variables, and the causal relationships between latent variables are further identified. This study theoretically proves that the proposed algorithm can correctly extract single-factor sets and further identify the causal structure among latent variables. Simulation data and real-world data experimental results further verify the correctness and effectiveness of the proposed algorithm.