Anytime-Valid Detection of Unauthorized Data Use in Authentication and Machine Learning Systems

Loading...

Date

2026

Journal Title

Journal ISSN

Volume Title

Attention Stats

Abstract

Unauthorized data use---whether through credential-database compromise or the incorporation of datasets into machine learning (ML) models' training without data owners' consent---poses a persistent threat to security, privacy, and data sovereignty. Detecting such unauthorized data use in computing systems requires mechanisms that remain effective and reliable even against adaptive adversaries. This dissertation develops anytime-valid detection frameworks for identifying unauthorized data use in authentication and machine learning systems, which provide tunable and provably bounded false-detection guarantees that remain valid under continuous evidence accumulation and adaptive stopping.

The first part of this dissertation studies credential-database breach detection using honeywords (also known as decoy passwords). We begin by analyzing the security of existing honeyword-based methods under realistic threat models where breach attackers exploit cross-site password knowledge and false-alarm attackers attempt to trigger breach alarms without breaching the credential database. Our analysis shows that state-of-the-art schemes fail to achieve near-ideal false-negative rates and are prone to false alarms, limiting their practical deployment. Building on these insights, we introduce LeakSentinel, a honeyword framework for anytime-valid detection of credential-database breaches. LeakSentinel guarantees a tunable and provably bounded global false-detection rate, meaning that in the absence of a breach, the probability of ever raising an alarm over an infinite sequence of login attempts is tunable and provably bounded. At the same time, it maintains strong detection power against adaptive breach attack strategies.

The second part addresses unauthorized data use in ML systems. We propose a general data-use auditing framework that enables a data owner to test whether her dataset has been used in model training. The framework achieves anytime-valid detection in that the data owner may continuously accumulate data-use evidence through querying the audited ML model and adaptively stop at any time while maintaining a tunable, provably bounded false-detection rate. Finally, we explore the limits of data-use auditing in adversarial scenarios through AcidWash, a framework for purifying auditable training data to mitigate data-use auditing in ML models. AcidWash demonstrates how adversaries can mitigate detection while preserving model utility, providing a useful tool to evaluate the robustness of data-use auditing methods.

Together, these contributions establish a principled foundation for anytime-valid detection of unauthorized data use and clarify the capabilities of statistical auditing in security-critical systems.

Description

Provenance

Subjects

Computer science, Anytime-valid detection, Authentication, Honeywords, Machine learning, Provably bounded false-detection rate, Unauthorized data-use detection

Citation

Citation

Huang, Zonghao (2026). Anytime-Valid Detection of Unauthorized Data Use in Authentication and Machine Learning Systems. Dissertation, Duke University. Retrieved from https://hdl.handle.net/10161/35318.

Collections


Except where otherwise noted, student scholarship that was shared on DukeSpace after 2009 is made available to the public under a Creative Commons Attribution / Non-commercial / No derivatives (CC-BY-NC-ND) license. All rights in student work shared on DukeSpace before 2009 remain with the author and/or their designee, whose permission may be required for reuse.