Biases in Association Studies and Heritability Estimation for Structured Populations

Limited Access
This item is unavailable until:
2027-05-06

Date

2026

Journal Title

Journal ISSN

Volume Title

Attention Stats

Abstract

Population structure and genetic relatedness are pervasive features of modern genomic studies and pose fundamental challenges for statistical inference in human genetics. Many modern methods for genetic association testing and heritability estimation rely on kinship matrices, also known as genetic relatedness matrices, to model correlations between individuals. However, common kinship estimators are now known to exhibit substantial biases in structured populations, raising concerns about the validity of downstream analyses.

This dissertation investigates how kinship estimation bias affects different layers of genetic inference and develops a unified framework that clarifies when such biases are inconsequential, when they induce systematic error, and how they can be addressed. First, it shows both theoretically and empirically that genetic association testing using principal component analysis and linear mixed-effects models is remarkably robust to a broad class of kinship estimation biases. Association test statistics are invariant to these biases because distortions in kinship are absorbed by nuisance parameters, an insight that extends to generalized linear models. This result provides reassurance regarding the validity of existing genome-wide association studies in structured populations.

In contrast, this dissertation demonstrates that heritability estimation behaves fundamentally differently and is intrinsically sensitive to kinship bias. Because variance component models depend directly on the scale of the kinship matrix, biased kinship estimation propagates into systematic bias in heritability estimates. Through analytical characterization and simulation-based investigation, this work clarifies how population structure, allele frequency spectrum, and estimator choice interact to distort heritability inference, establishing the necessity of unbiased and structure-aware kinship estimation in heterogeneous populations.

Finally, this dissertation addresses the downstream consequences of population structure for polygenic risk prediction. It develops a principled framework that connects linear mixed model–based genome-wide association studies with PRS construction by explicitly modeling how population structure and relatedness induce covariance between genetic loci. This framework yields structure-aware linkage disequilibrium estimation and a corresponding summary statistic representation, together with an effective sample size that quantifies information loss due to relatedness. These results enable existing PRS methods to be applied reliably in diverse and admixed populations without relying on discrete ancestry stratification.

Together, this work delineates the boundaries of robustness and vulnerability across key genetic inference tasks and provides a coherent statistical foundation for association testing, heritability estimation, and polygenic prediction in increasingly diverse genomic datasets. By clarifying the role of population structure across these settings, this dissertation contributes toward more reliable and equitable applications of genomic analysis in research and precision medicine.

Description

Provenance

Subjects

Biostatistics

Citation

Citation

Hou, Zhuoran (2026). Biases in Association Studies and Heritability Estimation for Structured Populations. Dissertation, Duke University. Retrieved from https://hdl.handle.net/10161/35123.

Collections


Except where otherwise noted, student scholarship that was shared on DukeSpace after 2009 is made available to the public under a Creative Commons Attribution / Non-commercial / No derivatives (CC-BY-NC-ND) license. All rights in student work shared on DukeSpace before 2009 remain with the author and/or their designee, whose permission may be required for reuse.