Developing Intelligent Reinforcement Learning Bots for Radiation Therapy Treatment Planning
Date
2026
Authors
Advisors
Journal Title
Journal ISSN
Volume Title
Attention Stats
Abstract
Radiation therapy (RT) is a cornerstone of modern cancer treatment, requiring precise dose delivery to tumors while minimizing exposure to surrounding healthy tissues. However, the manual design of such treatment (referred as treatment planning) is time-intensive and highly dependent on clinical expertise, thus likely subject to inter-planner variability. This dissertation aims to address these challenges by developing a human-AI collaborative framework that leverages reinforcement learning (RL) algorithms to automate RT treatment planning. The proposed work focuses on integrating clinical priorities and preferences into AI-driven workflows to enhance planning efficiency, consistency, and adaptability across diverse clinical scenarios.This dissertation is structured around three specific aims. Aim 1 focuses on the development of a multi-agent RL system for fluence map editing for whole-breast radiation therapy (WBRT) using the electronic compensation (ECOMP) technique. In this framework, each pixel in the fluence map was assigned an RL agent to perform independent actions. Projected dose profiles in beam’s-eye-view were generated as state inputs to the RL network. Pixel-wise actions were selected by predicting Q values to modify fluence maps and improve overall plan quality. After dose calculation, reward signals were calculated from variations in target coverage and dose homogeneity and fed back to update network parameters. The developed framework generated breast ECOMP plans within 2 minutes, achieving clinically comparable isodose distributions and dosimetric endpoints. Aim 2 extended this framework to head-and-neck cancer intensity-modulated radiation therapy (IMRT), addressing complex trade-offs between target coverage and organ-at-risk (OAR) sparing. A deep reinforcement learning (DRL) agent was developed to directly interact with a clinical treatment planning system to perform inverse optimization. During each optimization process, intermediate dosimetric endpoints, dose–volume constraint values, and structure objective function losses were collected as state information. By adjusting objective constraints as actions, the agent learned to optimize rewards that balanced planning target volume (PTV) coverage and OAR sparing. The trained agent was able to generate end-to-end treatment plans within the clinical platform, with an average planning time of 12.4±3.1 minutes and without human intervention. In addition to clinically comparable PTV coverage and improved OAR sparing, DRL-generated plans demonstrated reduced variability and lower monitor units (MU), indicating a more consistent planning strategy compared to manual planning. Parotid-sparing preferences were further encoded into the state space, enabling the agent to autonomously adapt sparing strategies based on clinical priorities. This framework efficiently generated simultaneous integrated boost (SIB) treatment plans tailored to varying clinical preferences. Aim 3 sought to further enhance adaptability through model predictive control for real-time strategy adjustment and clinical decision support. A two-stage approach was proposed. First, a Deep Dose Predictive (DDP) model was trained to predict dose responses based on historical plan states and dose–volume objective (DVO) adjustments, using datasets generated via Monte Carlo sampling without human intervention. In the second stage, the trained DDP model was employed to guide automatic DVO adjustments through model predictive control. By forecasting future plan states and evaluating predicted dose responses using a weighted score function, the system selected optimal adjustments aligned with specific clinical priorities without requiring model retraining. Validation in head-and-neck IMRT demonstrated clinically comparable plan quality with improved efficiency and adaptability. Overall, this work introduces a transformative framework for intelligent and flexible automation of radiation therapy treatment planning.
Type
Department
Description
Provenance
Subjects
Citation
Permalink
Citation
Yang, Dongrong (2026). Developing Intelligent Reinforcement Learning Bots for Radiation Therapy Treatment Planning. Dissertation, Duke University. Retrieved from https://hdl.handle.net/10161/35149.
Collections
Except where otherwise noted, student scholarship that was shared on DukeSpace after 2009 is made available to the public under a Creative Commons Attribution / Non-commercial / No derivatives (CC-BY-NC-ND) license. All rights in student work shared on DukeSpace before 2009 remain with the author and/or their designee, whose permission may be required for reuse.
