Creation and Optimization of Functional Annotation Models to Identify Gene Regulatory Effects of Genetic Variation.
Date
2026
Authors
Advisors
Journal Title
Journal ISSN
Volume Title
Attention Stats
Abstract
Despite a constantly increasing number of efforts to understand the function of all parts of the human genome, the exact functions and mechanisms caused by human genomic variation remain largely unknown. In genome-wide studies to assess the correlations between traits and genomic variation, each trait will have many correlated variants which will all be inherited together. The ability to determine potentially causal variation is made even more difficult by the high percentage of implicated variants that fall in non-coding regions of the genome; the function of these regions is much more subtle and difficult to observe. The major goal of the work presented in this dissertation is to address this issue by functionally annotating regions of the genome that have previously been correlated with molecular traits. In working towards that goal, computational methods have been developed which can better represent and analyze genomic data. One of these computational methods is a heuristic pipeline to improve alignment of genomic samples to a reference genome. Specifically, this method addresses the issue of alignment bias, where two different alleles of the same genomic variant will align to a reference genome at different rates depending on their similarities to the reference genome. Many assays which investigate variant effects rely on differential expression of alleles, and reference bias can be a major confounder in these experiments. After completing development of both this pipeline and a simulator for genomic data, I was able to test these methods. I showed that the rate of correct alignment of genomic data increased with the novel alignment approach over a standard alignment approach, and statistics indicating the quality of the alignment improved as well. As a step towards functional annotation of genomic regions, I have analyzed a set of STARR-seq data, which can provide an indication of regulatory activity of individual variants. Using this dataset, I assessed the relationship between STARR-seq data and functional assays, providing a potential mechanism for the effects found by the STARR-seq assay. In the pursuit of estimating accurate allelic effects of genetic variants, I expanded upon a previously published Bayesian graphical model for variant effect estimation, BIRD. Specifically, I developed an extension of this model which can take as input the results of multiple experiments to measure variant effects that were performed on different genomes. Due to this experimental design, genomic datasets can increase in size while avoiding dropout of alleles. Additionally, simulated variant effects were found to be more easily estimated with the aggregation of data across multiple experiments. Finally, I overlapped regions with regulatory activity identified by STARR-seq and regions which were correlated with traits. This analysis provided a set of functionally annotated regions by variant effects estimated from STARR-seq data, which had previously been associated with molecular traits. Further analysis of these data types by similar methods could be impactful for the identification of causes of deleterious traits and diseases.
Type
Department
Description
Provenance
Subjects
Citation
Permalink
Citation
Dura, Katherine (2026). Creation and Optimization of Functional Annotation Models to Identify Gene Regulatory Effects of Genetic Variation. Dissertation, Duke University. Retrieved from https://hdl.handle.net/10161/35231.
Collections
Except where otherwise noted, student scholarship that was shared on DukeSpace after 2009 is made available to the public under a Creative Commons Attribution / Non-commercial / No derivatives (CC-BY-NC-ND) license. All rights in student work shared on DukeSpace before 2009 remain with the author and/or their designee, whose permission may be required for reuse.
