A Framework for Automatic Failure Mining and Continuous Improvement of Healthcare AI Systems
Date
2026
Authors
Advisors
Journal Title
Journal ISSN
Volume Title
Attention Stats
Abstract
Healthcare AI systems deployed in real-world business-to-business settings often lack the direct feedback signals that support continuous improvement in consumer-facing applications. In particular, clinical AI services must operate under privacy, regulatory, and annotation-cost constraints, making it difficult to identify useful failures and update models reliably after deployment. This thesis proposes an Evolving-as-a-Service (EaaS) framework for the automatic mining of high-value failures and the continuous improvement of healthcare AI systems under limited-feedback conditions. The framework is studied in the setting of extractive clinical question answering over de-identified electronic health records using the emrQA-msquad dataset. A single-round offline simulation of deployment is constructed in which a baseline model processes online-like traffic, produces candidate answer sets, and is evaluated by a no-gold mining scorer designed to identify cases where the deployed top-1 prediction is likely suboptimal despite the presence of a stronger candidate. Mined cases are then converted into post-training data and used to compare two update strategies: pairwise preference-style training on mined outputs and supervised fine-tuning on gold spans from mined-selected examples. Experiments show that the proposed mining loop can recover useful bad cases from simulated online traffic without access to gold labels during the mining stage. While pairwise pseudo-label training fails to yield reliable held-out gains, supervised fine-tuning on gold spans from mined-selected cases produces consistent improvements across multiple model families. On held-out validation, BERT improves from 0.0077 to 0.0482 in exact match and from 0.1450 to 0.3850 in token-level F1; RoBERTa improves from 0.0128 to 0.0608 in exact match and from 0.2975 to 0.4640 in token-level F1; and Qwen improves from 0.0720 to 0.1530 in exact match and from 0.1562 to 0.3504 in token-level F1. These results show that mined online failures are valuable primarily as a data selection mechanism rather than as fully reliable pseudo-labels. More broadly, the thesis reframes post-deployment model improvement as a service problem with explicit stages for logging, mining, curation, post-training, and protected evaluation. The proposed framework offers a practical foundation for continuous improvement in healthcare AI systems, while highlighting that future progress depends on stronger mining signals, better long-context handling, and careful governance for real-world deployment.
Type
Description
Provenance
Subjects
Citation
Permalink
Citation
Li, Tianhao (2026). A Framework for Automatic Failure Mining and Continuous Improvement of Healthcare AI Systems. Master's thesis, Duke University. Retrieved from https://hdl.handle.net/10161/35074.
Collections
Except where otherwise noted, student scholarship that was shared on DukeSpace after 2009 is made available to the public under a Creative Commons Attribution / Non-commercial / No derivatives (CC-BY-NC-ND) license. All rights in student work shared on DukeSpace before 2009 remain with the author and/or their designee, whose permission may be required for reuse.
