Design of Safer Prodrugs Powered by Machine Learning

Limited Access
This item is unavailable until:
2028-06-06

Date

2026

Journal Title

Journal ISSN

Volume Title

Repository Usage Stats

1
views
0
downloads

Attention Stats

Abstract

Prodrugs are easily deployable chemical entities with beneficial pharmacokinetic properties that harness biotransformations to generate active pharmaceutical ingredients (APIs) in situ. However, the rational design of prodrugs is challenging as it requires careful crafting of release mechanisms and holistic optimization of pharmacokinetic properties. Accordingly, prodrug design often relies on simple functionalizations while more complex prodrugs have typically been discovered serendipitously. Rational design of more advanced prodrugs could enable the delivery of life-saving medications currently inaccessible by prodrug design with improved safety and efficacy. Machine learning is poised to support rational design of prodrugs through structural-based predictions to guide prodrug development towards ideal absorption, distribution, metabolism, excretion, and toxicity (ADMET) profiles, ease of synthesis, non-toxic metabolite formation, and minimal activity of the inactive prodrug. Coupled with generative models, millions of systematically proposed structures can be efficiently filtered to the most promising candidates for in vitro and in vivo validation. In this pursuit, we designed and validated a novel machine learning pipeline for rapid and systematic design of prodrugs with desired pharmacokinetic properties through the following:(1) We holistically characterized the current landscape of all FDA-approved small molecule prodrugs and found that most prodrugs are simple while more complex prodrugs that can reduce side effects have been discovered serendipitously. This highlighted the need for advanced computational workflows to derisk and accelerate the identification of novel prodrugs; (2) We conceived, implemented, and validated a novel pairwise machine learning approach, DeepDelta, to directly predict ADMET improvements of molecular derivatizations. This approach significantly improved the performance of neural network and tree-based machine learning models across 10 ADMET benchmark tasks. By training on the direct comparison of two molecules and providing combinatorically expanded amounts of data to learn from, our approach provides models increased predictive resolution for the design of prodrug derivatives that solve specific ADMET liabilities of existing drugs; (3) We conceived, implemented, and validated ActiveDelta, an adaptive approach that leverages paired molecular representations to predict improvements from the current best training compound to prioritize further data acquisition. In low data regimes, both neural network and tree-based machine learning models using the ActiveDelta approach selected more of the most potent leads than any standard single-molecule active learning approach across 99 benchmarking and external datasets. ActiveDelta implementations of deep models also not only found the most chemically diverse hits in terms of their unique scaffolds, with potential to create multiple lead series to enable further development, but also enriched the scaffold diversity of “negative” training data to improve future compound selection; (4) We conceived, implemented, and validated DeltaClassifier, a novel molecular pre-processing and pairing approach that incorporates bounded data, counteracting skewed class proportions and providing valuable chemical diversity during training. Both neural network and tree-based machine learning models using the DeltaClassifier outperform several state-of-the-art molecular regression models at correctly anticipating which of two molecules are expected to be more potent. The improvement of DeltaClassifiers over traditional methods correlates with the amount of bounded data in the training data, suggesting that this concept is particularly useful for datasets and medicinal chemistry optimization in early project stages where data is incompletely characterized or limited; (5) We coupled these predictive machine learning approaches with generative models to a create machine learning-driven platform for the rapid and systematic design of prodrugs that we validated with two case studies. First, we generated alternatives to the FDA-approved antibiotic prodrug cefuroxime axetil that showed improved ex vivo bioavailability and reduced microbiome disruption in vitro compared to the currently marketed prodrug. Second, we created first-in-class prodrugs of the Bcl-xL inhibitor A-1331852, a compound with no existing prodrug formulations, which retained anti-cancer efficacy while reducing platelet toxicity in vivo; Overall, this dissertation describes the creation of an experimentally validated machine learning pipeline to enable computational prodrug design with higher resolution and provides novel prodrug formulations for multiple indications with improved safety profiles.

Description

Provenance

Subjects

Computational chemistry, Pharmacology

Citation

Citation

Fralish, Zachary (2026). Design of Safer Prodrugs Powered by Machine Learning. Dissertation, Duke University. Retrieved from https://hdl.handle.net/10161/35107.

Collections


Except where otherwise noted, student scholarship that was shared on DukeSpace after 2009 is made available to the public under a Creative Commons Attribution / Non-commercial / No derivatives (CC-BY-NC-ND) license. All rights in student work shared on DukeSpace before 2009 remain with the author and/or their designee, whose permission may be required for reuse.