A Framework for Automatic Failure Mining and Continuous Improvement of Healthcare AI Systems
| dc.contributor.advisor | Yang, Zhenyu | |
| dc.contributor.author | Li, Tianhao | |
| dc.date.accessioned | 2026-07-06T19:50:14Z | |
| dc.date.issued | 2026 | |
| dc.department | DKU - Medical Physics Master of Science Program | |
| dc.description.abstract | Healthcare AI systems deployed in real-world business-to-business settings often lack the direct feedback signals that support continuous improvement in consumer-facing applications. In particular, clinical AI services must operate under privacy, regulatory, and annotation-cost constraints, making it difficult to identify useful failures and update models reliably after deployment. This thesis proposes an Evolving-as-a-Service (EaaS) framework for the automatic mining of high-value failures and the continuous improvement of healthcare AI systems under limited-feedback conditions. The framework is studied in the setting of extractive clinical question answering over de-identified electronic health records using the emrQA-msquad dataset. A single-round offline simulation of deployment is constructed in which a baseline model processes online-like traffic, produces candidate answer sets, and is evaluated by a no-gold mining scorer designed to identify cases where the deployed top-1 prediction is likely suboptimal despite the presence of a stronger candidate. Mined cases are then converted into post-training data and used to compare two update strategies: pairwise preference-style training on mined outputs and supervised fine-tuning on gold spans from mined-selected examples. Experiments show that the proposed mining loop can recover useful bad cases from simulated online traffic without access to gold labels during the mining stage. While pairwise pseudo-label training fails to yield reliable held-out gains, supervised fine-tuning on gold spans from mined-selected cases produces consistent improvements across multiple model families. On held-out validation, BERT improves from 0.0077 to 0.0482 in exact match and from 0.1450 to 0.3850 in token-level F1; RoBERTa improves from 0.0128 to 0.0608 in exact match and from 0.2975 to 0.4640 in token-level F1; and Qwen improves from 0.0720 to 0.1530 in exact match and from 0.1562 to 0.3504 in token-level F1. These results show that mined online failures are valuable primarily as a data selection mechanism rather than as fully reliable pseudo-labels. More broadly, the thesis reframes post-deployment model improvement as a service problem with explicit stages for logging, mining, curation, post-training, and protected evaluation. The proposed framework offers a practical foundation for continuous improvement in healthcare AI systems, while highlighting that future progress depends on stronger mining signals, better long-context handling, and careful governance for real-world deployment. | |
| dc.identifier.uri | ||
| dc.rights.uri | ||
| dc.subject | Computer science | |
| dc.title | A Framework for Automatic Failure Mining and Continuous Improvement of Healthcare AI Systems | |
| dc.type | Master's thesis | |
| duke.embargo.months | 23 | |
| duke.embargo.release | 2028-06-06T19:50:14Z |