Modeling, Analysis, and Interpretation of High-Throughput Genomic Perturbation Assays

Loading...

Date

2026

Journal Title

Journal ISSN

Volume Title

Attention Stats

Abstract

High-throughput genomic perturbation assays provide large-scale measurements of molecular response to targeted genetic interventions, but the resulting data are typically sparse, noisy, and characterized by rare or weak signals. In this regime, biological ambition outpaces naïve statistical practice: detecting meaningful regulatory effects requires procedures that remain calibrated under dilution, heterogeneity, and limited signal strength. This dissertation develops statistical foundations for calibrated inference in such perturbation systems. We first introduce a generative framework for gRNA-level count data from single-cell CRISPR assays that represents observed counts as the superposition of signal and nuisance components, each governed by an occurrence mechanism and a conditional magnitude distribution. This concurrent-process formulation reproduces key aggregate features of empirical UMI count distributions and provides a structured statistical representation of sparsity and noise in the data-generating process. Building on this foundation, we develop FRACTEL, a bounded-rank aggregation procedure for combining multiple gRNA-level tests targeting a shared regulatory element. Rather than diffusing evidence across all guides or relying solely on the single most extreme result, FRACTEL evaluates the least likely among a bounded set of leading order statistics, concentrating evidentiary weight where signal is most plausibly localized while controlling the depth of aggregation. Empirical analyses characterize how power depends on signal architecture, guide multiplicity, and baseline expression, illustrating the impact of dilution and heterogeneous activity on detection.

To enable principled calibration beyond empirical resampling, we derive exact and asymptotic distributional results for finite-rank order-statistic aggregation statistics. We obtain explicit null distributions and critical values in finite samples, establish fixed-rank asymptotics when the aggregation depth remains small relative to the number of tests, and show how evidentiary weight is distributed non-uniformly across ordered p-values as the rank bound varies. These results provide calibrated inference in small-sample and truncated regimes and clarify the operating characteristics of bounded-rank testing procedures. Finally, we apply rank-based and allele-controlled statistical analyses to in silico saturated mutagenesis experiments using a predictive regulatory model. By constructing local perturbation landscapes in constitutively closed genomic regions and in candidate cis-regulatory elements, we identify context-dependent directional asymmetries consistent with regulatory constraint: variants observed in the human population are relatively depleted among mutations predicted to increase activity in closed regions, whereas in active regulatory elements they are relatively depleted among mutations predicted to decrease activity. Together, these findings link calibrated statistical methodology to biologically interpretable patterns of regulatory constraint in high-throughput perturbation experiments.

Description

Provenance

Subjects

Statistics, Bioinformatics, genomic perturbation, rank aggregation, single-cell analysis, statistical inference

Citation

Citation

Doty, Richard (2026). Modeling, Analysis, and Interpretation of High-Throughput Genomic Perturbation Assays. Dissertation, Duke University. Retrieved from https://hdl.handle.net/10161/35330.

Collections


Except where otherwise noted, student scholarship that was shared on DukeSpace after 2009 is made available to the public under a Creative Commons Attribution / Non-commercial / No derivatives (CC-BY-NC-ND) license. All rights in student work shared on DukeSpace before 2009 remain with the author and/or their designee, whose permission may be required for reuse.