Domain-centric Methods for Robust Speech Enhancement Models in Cochlear Implants in Dynamic Acoustic Environments
Date
2026
Authors
Advisors
Journal Title
Journal ISSN
Volume Title
Repository Usage Stats
views
downloads
Attention Stats
Abstract
Hearing loss affects millions of individuals worldwide and significantly impairs communication and quality of life. Cochlear implants (CIs) provide auditory perception to individuals with severe-to-profound hearing loss; however, CI users often experience substantial difficulty understanding speech in noisy and reverberant environments. Although numerous speech enhancement strategies have been developed for automatic speech processing applications, several key challenges limit their applicability to CI systems, including poor generalization to unseen acoustic environments, high computational complexity incompatible with CI processors, and the limited representational capacity of magnitude-only masking strategies that ignore phase information. Addressing these challenges requires the development of computationally efficient and robust speech enhancement approaches specifically designed for the constraints of CI signal processing.
This dissertation investigates machine learning-based speech enhancement strategies that leverage phonetic structure and domain-robust training techniques to improve speech intelligibility for CI users in adverse listening environments. First, the effects of acoustic domain mismatch arising from unseen combinations of noise and reverberation are systematically analyzed, and training strategies that promote robustness across diverse acoustic conditions are identified. Second, a computationally efficient alternative to mixture-of-experts (MoE) architectures is proposed, replacing multiple phoneme-specific experts with a single-expert model while maintaining phoneme-dependent enhancement capabilities. The impact of broad phonetic grouping strategies on speech enhancement performance in noisy and reverberant conditions is further examined to reduce phoneme classification errors.
To improve mask estimation quality, phase-sensitive masks (PSMs) are investigated as training targets to incorporate phase information neglected by conventional magnitude-based masks. The influence of different output activation functions and pre-emphasis filtering on PSM estimation stability and performance is evaluated. Finally, a lightweight adversarial training framework is introduced to improve phonetic classification robustness under acoustic domain mismatch. By integrating adapter modules into a recurrent neural network architecture and employing a domain-adversarial training strategy, the proposed method learns environment-invariant phonetic representations while maintaining computational efficiency suitable for CI deployment.
Experimental results demonstrate that the proposed approaches improve speech enhancement robustness and intelligibility for CI users across a wide range of noisy and reverberant conditions. Collectively, these findings provide new insights into phoneme-aware speech enhancement, domain-robust training strategies, and computationally efficient neural architectures for auditory prostheses, contributing toward the development of practical machine learning-based front-end processing strategies for future CI systems.
Type
Department
Description
Provenance
Subjects
Citation
Permalink
Citation
Saha, Sohini (2026). Domain-centric Methods for Robust Speech Enhancement Models in Cochlear Implants in Dynamic Acoustic Environments. Dissertation, Duke University. Retrieved from https://hdl.handle.net/10161/35277.
Collections
Except where otherwise noted, student scholarship that was shared on DukeSpace after 2009 is made available to the public under a Creative Commons Attribution / Non-commercial / No derivatives (CC-BY-NC-ND) license. All rights in student work shared on DukeSpace before 2009 remain with the author and/or their designee, whose permission may be required for reuse.
