A Framework for Automatic Failure Mining and Continuous Improvement of Healthcare AI Systems

Limited Access
This item is unavailable until:
2028-06-06

Date

2026

Journal Title

Journal ISSN

Volume Title

Attention Stats

Abstract

Healthcare AI systems deployed in real-world business-to-business settings often lack the direct feedback signals that support continuous improvement in consumer-facing applications. In particular, clinical AI services must operate under privacy, regulatory, and annotation-cost constraints, making it difficult to identify useful failures and update models reliably after deployment. This thesis proposes an Evolving-as-a-Service (EaaS) framework for the automatic mining of high-value failures and the continuous improvement of healthcare AI systems under limited-feedback conditions. The framework is studied in the setting of extractive clinical question answering over de-identified electronic health records using the emrQA-msquad dataset. A single-round offline simulation of deployment is constructed in which a baseline model processes online-like traffic, produces candidate answer sets, and is evaluated by a no-gold mining scorer designed to identify cases where the deployed top-1 prediction is likely suboptimal despite the presence of a stronger candidate. Mined cases are then converted into post-training data and used to compare two update strategies: pairwise preference-style training on mined outputs and supervised fine-tuning on gold spans from mined-selected examples. Experiments show that the proposed mining loop can recover useful bad cases from simulated online traffic without access to gold labels during the mining stage. While pairwise pseudo-label training fails to yield reliable held-out gains, supervised fine-tuning on gold spans from mined-selected cases produces consistent improvements across multiple model families. On held-out validation, BERT improves from 0.0077 to 0.0482 in exact match and from 0.1450 to 0.3850 in token-level F1; RoBERTa improves from 0.0128 to 0.0608 in exact match and from 0.2975 to 0.4640 in token-level F1; and Qwen improves from 0.0720 to 0.1530 in exact match and from 0.1562 to 0.3504 in token-level F1. These results show that mined online failures are valuable primarily as a data selection mechanism rather than as fully reliable pseudo-labels. More broadly, the thesis reframes post-deployment model improvement as a service problem with explicit stages for logging, mining, curation, post-training, and protected evaluation. The proposed framework offers a practical foundation for continuous improvement in healthcare AI systems, while highlighting that future progress depends on stronger mining signals, better long-context handling, and careful governance for real-world deployment.

Description

Provenance

Subjects

Computer science

Citation

Citation

Li, Tianhao (2026). A Framework for Automatic Failure Mining and Continuous Improvement of Healthcare AI Systems. Master's thesis, Duke University. Retrieved from https://hdl.handle.net/10161/35074.

Collections


Except where otherwise noted, student scholarship that was shared on DukeSpace after 2009 is made available to the public under a Creative Commons Attribution / Non-commercial / No derivatives (CC-BY-NC-ND) license. All rights in student work shared on DukeSpace before 2009 remain with the author and/or their designee, whose permission may be required for reuse.