Latent Variable Modeling and Analysis of Vocalizations
Date
2026
Authors
Advisors
Journal Title
Journal ISSN
Volume Title
Repository Usage Stats
views
downloads
Attention Stats
Abstract
Vocal communication is the basis of complex interactions in both humans and other animals. Understanding the structure of vocalization provides insight into learning, socialization, and communication, in addition to the neural activity supporting these behaviors. However, quantifying vocal production is difficult: what is the ``right'' information to represent in a bird's song, a gerbil's call, or an infant's babbling? Methods aimed at answering these questions need to capture both fine-timescale variability, on the order of milliseconds, and long-timescale variability, on the order of weeks. Only by capturing both of these attributes can we understand typical vocal variability, differences in vocal production in social contexts, and changes due to learning.
Latent variable models offer a promising approach for investigating complex vocalizations. These models learn a reduced set of variables that are easier to analyze than the audio itself while still reflecting the structure of the original data. However, the variables learned by these models are difficult to visualize and interpret. In the case of vocal production, these latent variables have complex and opaque relationships with the physical mechanisms of vocal control that we wish to model.
In this work, I develop and apply methods for computational neuroethology: using computational modeling to link the brain and behavior in animals. I begin with a general introduction to common latent variable models used in neuroscience for analysis and visualization. I then develop three latent variable models aimed at addressing the shortcomings of current models. The first two, presented in chapter two, address issues of reproducibility and interpretability of current general-purpose latent variable models. I develop the Rosetta-VAE, a method for aligning latent spaces across variational autoencoders (VAEs), demonstrating that using just a small set of reference points we can train variational autoencoders with closely matched latent representations. I then create Quasi-Monte Carlo Latent Variable Models (QLVMs), a new class of latent variable models that can be directly visualized and analyzed. I show that these models use their latent spaces more efficiently than VAEs and that their low dimensionality allows for direct analysis and visualization of their latent variables.In chapter 3, I develop a latent variable model grounded in biophysical models of vocalization, specifically designed for studying vocal control. I demonstrate that this model can accurately model vocalizations from multiple species, that the learned representations can be used to assess vocal similarity and track change over time, and that the latent parameters have close links to physiology. Lastly, in chapter 4, I cover my work studying vocal learning in zebra finches. I demonstrate that zebra finches learn to sing using dopaminergic reinforcement --- the same mechanism as in learning externally reinforced behaviors. I additionally show that information about vocal production is represented strongly and consistently by calcium activity in song learning circuitry. The tools developed in this dissertation enable more effective study of complex, naturalistic vocal control, offering insight into the neural and social mechanisms of learning and communication.
Type
Department
Description
Provenance
Subjects
Citation
Permalink
Citation
Martinez, Miles (2026). Latent Variable Modeling and Analysis of Vocalizations. Dissertation, Duke University. Retrieved from https://hdl.handle.net/10161/35236.
Collections
Except where otherwise noted, student scholarship that was shared on DukeSpace after 2009 is made available to the public under a Creative Commons Attribution / Non-commercial / No derivatives (CC-BY-NC-ND) license. All rights in student work shared on DukeSpace before 2009 remain with the author and/or their designee, whose permission may be required for reuse.
