Research

My research develops statistical foundations and methodology for learning interpretable latent structure from complex data. A recurring goal is to understand when latent structures are identifiable, how they can be learned and quantified reliably in high dimensions, and how statistically grounded representations can support scientific measurement and modern AI.

For a complete chronological list of papers and preprints, see Publications.

In the selected work below, underlined names are students or postdoctoral researchers under my supervision; ✉ marks corresponding authors, * marks co-first authors, and ⁺ marks alphabetical authorship.

Recurring questions

  • Identifiability. Which features of a latent representation or architecture are determined by the distribution of the observed data, and under what structural assumptions?
  • Learning and inference. How can latent structure be recovered efficiently in high dimensions, with finite-sample guarantees and principled uncertainty quantification?
  • Measurement and interpretation. How can latent representations support interpretable measurement of heterogeneous populations, scientific constructs, and complex AI systems?

These questions recur across three connected research directions: (a) Identifiable deep and causal representations; (b) High-dimensional statistical inference with latent structure, and (c) Statistical measurement for humans and AI systems.

Research overview showing observed data, structured latent representations, identifiable causal structure, statistical inference, and interpretable measurement.

1. Identifiable deep and causal representations

I study how latent representations in deep generative models can be made structurally meaningful rather than arbitrary. This work develops identifiable models with discrete latent layers, including Bayesian Pyramids and Deep Discrete Encoders, and investigates when latent dimension, graphical structure, and other features of a latent architecture can be recovered from observed data.

A related line of work studies causal structure among latent variables. I develop theory and methods for discrete causal representation learning, identifiability of latent causal graphical models, and tensor-unfolding approaches to learning unknown latent bipartite architectures. The broader aim is to connect expressive representation models with assumptions that make their learned structure statistically interpretable.

Selected work

    • Discrete Causal Representation Learning
      Wenjin Zhang, Yixin Wang, and Yuqi Gu✉. arXiv preprint arXiv:2603.25017 (2026) [arXiv] [Code]
    • Deep Discrete Encoders: Identifiable Deep Generative Models for Rich Data with Discrete Latent Layers
      Seunghyun Lee, and Yuqi Gu✉
      Journal of the American Statistical Association (2026), 121 (553): 194–208. [arXiv] [Journal] [Code]
    • On Theoretical Identifiability of Discrete Latent Causal Graphical Models
      Seunghyun Lee, and Yuqi Gu
      Transactions on Machine Learning Research (2026) [arXiv]
    • Constructive Q-Matrix Identifiability via Novel Tensor Unfolding
      Yuqi Gu✉
      Psychometrika (2026), 91: 131–150. [Journal] [PDF]
    • Blessing of Dependence: Identifiability and Geometry of Discrete Models with Multiple Binary Latent Variables
      Yuqi Gu✉
      Bernoulli (2025), 31 (2): 948–972. [arXiv] [Journal]
    • Bayesian Pyramids: Identifiable Multilayer Discrete Latent Structure Models for Discrete Data
      Yuqi Gu✉, and David B. Dunson
      Journal of the Royal Statistical Society Series B: Statistical Methodology (2023), 85 (2): 399–426. [arXiv] [Journal] [Code]

2. High-dimensional statistical inference with latent structure

I develop scalable methods for high-dimensional data in which heterogeneity is governed by latent classes, mixed membership, factors, or clusters. The methods draw on spectral, tensor, likelihood-based, and geometric ideas, with an emphasis on finite-sample guarantees, optimal recovery, and uncertainty quantification.

Recent work moves beyond idealized local-independence assumptions by modeling latent graphical dependence, local dependence, and shared block structure. I also study transfer and adaptation across related populations. This program asks not only whether latent structure is identifiable, but also when computational procedures recover it at statistically optimal rates.

Selected work

    • Beyond Local Independence: High-Dimensional Latent Class Graphical Models with Shared Block Structure
      Seunghyun Lee, and Yuqi Gu✉. arXiv preprint arXiv:2606.29631 (2026) [arXiv]
    • Minimax-Optimal Spectral Clustering with Covariance Projection for High-Dimensional Anisotropic Mixtures
      Chengzhu Huang, and Yuqi Gu✉. arXiv preprint arXiv:2502.02580 (2025) [arXiv]
    • Generalized Grade-of-Membership Estimation for High-dimensional Locally Dependent Data
      Ling Chen*, Chengzhu Huang*, and Yuqi Gu✉
      Journal of the American Statistical Association (2026), accepted. [arXiv] [Journal] [Code]
    • Adaptive Transfer Clustering: A Unified Framework
      Yuqi Gu⁺, Zhongyuan Lyu⁺, and Kaizheng Wang⁺
      Journal of the American Statistical Association (2026), accepted. [arXiv] [Journal] [Code]
    • Spectral Clustering with Likelihood Refinement for High-dimensional Latent Class Recovery
      Zhongyuan Lyu, and Yuqi Gu✉
      Psychometrika (2026), 91: 695–723. [arXiv] [Journal]
    • Degree-heterogeneous Latent Class Analysis for High-dimensional Discrete Data
      Zhongyuan Lyu, Ling Chen, and Yuqi Gu✉
      Journal of the American Statistical Association (2025), 120 (552): 2435–2448. [arXiv] [Journal] [Code]

3. Statistical measurement for humans and AI systems

Psychometrics provides an intellectual foundation for my work on statistical measurement through latent-variable models. I develop models and theory for measuring latent attributes, response processes, and heterogeneous performance in educational, psychological, biomedical, and other scientific settings.

My recent work extends this measurement perspective to modern AI evaluation. Rather than reducing model behavior to a single aggregate score, I study structured variation in capabilities, reasoning behavior, semantic item structure, response accuracy, and chain-of-thought length. The goal is to build statistically grounded measurement frameworks for increasingly complex AI systems while maintaining a clear connection to the broader theory of latent-variable modeling.

Selected work

    • Scalable Text-Embedding-informed Cognitive Diagnosis of Large Language Models
      Jia Liu, Zhiyu Xu, and Yuqi Gu✉. arXiv preprint arXiv:2603.14676 (2026) [arXiv]
    • Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length
      Zhiyu Xu, Jia Liu, Yixin Wang, and Yuqi Gu✉
      Annals of Applied Statistics (2026), accepted. [arXiv] [Code]
    • A Blockwise Mixed Membership Model for Multivariate Longitudinal Data: Discovering Clinical Heterogeneity and Identifying Parkinson’s Disease Subtypes
      Kai Kang✉, and Yuqi Gu✉
      Annals of Applied Statistics (2026), 20 (1): 452–475. [arXiv] [Journal]
    • Deep Generative Modeling for Cognitive Diagnosis via Exploratory DeepCDMs
      Jia Liu, and Yuqi Gu✉
      Psychometrika (2026), 91: 151–176. [Journal] [PDF]
    • A Joint MLE Approach to Large-Scale Structured Latent Attribute Analysis
      Yuqi Gu✉, and Gongjun Xu
      Journal of the American Statistical Association (2023), 118 (541): 746–760. [arXiv] [Journal]
    • Dimension-Grouped Mixed Membership Models for Multivariate Categorical Data
      Yuqi Gu✉, Elena A. Erosheva, Gongjun Xu, and David B. Dunson
      Journal of Machine Learning Research (2023), 24 (88): 1–49. [arXiv] [Journal]

For the complete chronological record, including all older publications, see Publications.