High-dose-rate (HDR) prostate brachytherapy requires placement of multiple needles and programming of dwell times to achieve target coverage while sparing organs at risk. In current practice, needle placement and dwell-time selection depend heavily on physician experience. The authors investigated whether an automated approach using reinforcement learning (RL) could recommend needle positions and dwell times at the pre-planning stage to reduce procedure time and improve consistency of plan quality.
The RL agent observed an environment representing the patient anatomy and iteratively adjusted treatment parameters. Specifically, the agent was trained to select and modify the position of one chosen needle and set all dwell times on that needle to maximize a predefined reward function. After adjusting a needle, the agent proceeded to the next needle. This cycle continued until all needles were adjusted, and the agent could repeat multiple rounds up to a maximum number. The workflow therefore combined sequential needle-by-needle adjustments with repeated rounds of optimization driven by the RL policy.
Plan data from 100 prostate HDR boost patients treated in the authors’ clinic were included. Three RL models were trained separately using different amounts of training data: one trained on 5 patients, one on 10 patients, and one on 20 patients. An evaluation set of 5 patients was used as reported, and the remaining 75 patients comprised the test set. The source reports these cohort splits but does not provide additional demographic or clinical characteristics of the patients in the PubMed abstract.
The study compared dosimetry metrics and the number of needles used in the RL-generated plans against the clinical (ground truth) plans. Dosimetric endpoints highlighted in the report include Rectum D2cc, Urethra D20%, Prostate V150, and Prostate V100; reported comparisons were made after normalizing both RL and clinical plans to Prostate V100 = 95%. The number of needles used per plan was also recorded and compared between RL and clinical plans.
Across all three RL models trained with different numbers of patients, the average number of needles used in RL plans was 14, matching the clinical plans on average. When both plans were normalized to Prostate V100 = 95%, RL plans showed statistically significant improvements relative to clinical plans in several endpoints: a reduction in Rectum D2cc by 3.5%, a reduction in Urethra D20% by 2%, and a reduction in Prostate V150 by at least 5.5%. The three RL models demonstrated similar performance; there was no statistically significant difference among them for Rectum D2cc and Urethra D20% according to the information reported.
The authors present this work as the first study demonstrating that RL can autonomously generate clinically acceptable HDR prostate brachytherapy plans. Key advantages noted in the abstract include minimal data requirements for model training and apparent generalizability across different training set sizes. Because the RL agent produces anatomy-based needle positions and dwell times during pre-planning, the method has potential to shorten intraoperative procedure time, standardize plan quality, and reduce variability that arises from individual physician experience. The abstract does not report procedural time reductions in quantitative terms or provide details on computational runtime, user interaction, or integration with clinical planning systems.
The study concludes that an RL-based method can achieve equal or improved plan quality compared to conventional clinical approaches for HDR prostate brachytherapy and that the approach has substantial potential to standardize planning and reduce variability. Reported keywords include brachytherapy, high dose rate, prostate cancer, and reinforcement learning. MeSH terms listed in the source include Automation; Brachytherapy; Prostatic Neoplasms/radiotherapy; Radiation Dosage; Radiotherapy Dosage; Radiotherapy Planning, Computer-Assisted/methods; and Reinforcement Machine Learning.