---
title: "Deep Reinforcement Learning for Robust, Quality-of-Life–Aware NSCLC Treatment Protocols"
id: "biorxiv-12-robust-and-quality-of-life-aware-treatment-protocols-in-nsclc-using-deep"
canonical_url: "https://medichelpline.com/clinical-feed/biorxiv-12-robust-and-quality-of-life-aware-treatment-protocols-in-nsclc-using-deep"
content_type: "clinical_feed_article"
specialty: "Oncology"
source_name: "bioRxiv (Biomedical Preprints)"
source_url: "https://www.biorxiv.org/content/10.64898/2026.09.01.748242v1?rss=1"
published_at: "2026-09-04T12:00:00.000Z"
evidence_level: "Verified Feed"
license: "CC-BY-NC-4.0 / Informational Use"
---
# Deep Reinforcement Learning for Robust, Quality-of-Life–Aware NSCLC Treatment Protocols
## Provenance & Clinical Metadata
- **Canonical URL:** https://medichelpline.com/clinical-feed/biorxiv-12-robust-and-quality-of-life-aware-treatment-protocols-in-nsclc-using-deep
- **Specialty:** [Oncology](https://medichelpline.com/clinical-feed/oncology.md)
- **Primary Source:** bioRxiv (Biomedical Preprints)
- **Source URL:** [Original Journal Publication](https://www.biorxiv.org/content/10.64898/2026.09.01.748242v1?rss=1)
- **Published At:** 2026-09-04T12:00:00.000Z
- **Evidence Rating:** Verified Feed
## Executive GIST (TL;DR)
- The study applies a **deep reinforcement learning (DRL)** agent, informed by a two-population tumour growth model, to design treatment schedules for patients with **non-small cell lung cancer (NSCLC)**. - Virtual patients were generated using parameters previously fitted to clinical data from NSCLC patients treated with erlotinib; the agent was trained on this cohort. - Conventional systemic practice often uses **maximum tolerable dose (MTD)** until toxicity or progression, which can promote resistance; evolutionary therapy and adaptive therapy aim to delay resistance by exploiting eco-evolutionary dynamics. - The study compares three protocols: the DRL-derived policy, the adaptive therapy protocol of **Zhang et al.**, and **MTD**, using multiple metrics beyond time to progression (TTP). - Key performance metrics include **time to progression (TTP)**, a new robustness metric called **margin-to-failure (MTF)** that quantifies robustness to delayed treatment restart, and **quality-adjusted survival (QAS)** to capture patient QoL preferences. - The DRL policy produced higher median TTP, greater MTF, and improved QAS across all tested treatment decision intervals compared with the Zhang et al. protocol and MTD. - As decision intervals (time between dosing adjustments) increased, TTP under the DRL policy declined gradually toward the TTP achieved under MTD; the Zhang et al. protocol showed inconsistent performance and risked premature progression. - A population-level DRL policy trained on a cohort yielded an interpretable treatment rule that extended TTP for most previously unseen virtual patients and recommended resuming treatment at a lower tumour burden when monitoring is less frequent. - The study also explored **reward shaping** to encode different QoL preferences into the DRL objective, demonstrating how treatment strategies could be adjusted to reflect individual patient values. - The authors provide code/data links in a GitLab repository and declare no competing interests; funding sources include NWO and a Marie Skłodowska-Curie grant as reported.
## Clinical Analysis & Structured Key Points
Under current systemic treatment of metastatic cancer, a drug is frequently prescribed at maximum tolerable dose (MTD) until either unacceptable toxicity or progression. Unfortunately, in many patients this treatment strategy leads to the development of treatment resistance. Evolutionary therapy approaches aim to forestall or delay treatment resistance in cancer by exploiting eco-evolutionary interactions. A well-known implementation is the adaptive therapy protocol of Zhang et al., in which tumour burden thresholds are used to guide strategic treatment holidays. Deep reinforcement learning (DRL) has recently been used to optimise these approaches. However, research combining DRL with evolutionary therapy approaches has so far focused on time to progression (TTP) as a performance metric, and has not quantified safety in terms of robustness to delayed treatment restart or included patient preferences regarding quality of life (QoL) in treatment design. In our study, we use a DRL agent informed by a mathematical two-population tumour growth model to design treatment schedules for patients with non-small cell lung cancer (NSCLC). The agent is trained on a virtual patient cohort using parameters previously fitted to data from patients with NSCLC treated with erlotinib. Beyond TTP, we focus on improving robustness to delayed treatment restart and on how individual preferences and values impact QoL experienced during treatment. We compare TTP, robustness and QoL under the DRL policy, the adaptive therapy protocol of Zhang et al., and MTD. We introduce a robustness metric ``margin-to-failure'' (MTF), and compare quality-adjusted-survival (QAS) across different patient preference profiles. Finally, we explore reward shaping to assess how QoL preferences can be incorporated into DRL-based treatment design. To evaluate our results, we consider different decision intervals, defined as the time between dosing adjustments. The DRL policy achieved greater median TTP, MTF, and QAS across all treatment decision intervals compared to the other two protocols. As decision intervals increased, TTP under DRL declined gradually towards that achieved under MTD. In contrast, the Zhang et al. protocol performed inconsistently and could result in premature progression. Additionally, a population-level policy trained on a cohort of virtual patients produced an interpretable treatment rule that extended TTP for most previously unseen patients and indicated that treatment should resume at a lower tumour burden when monitoring is less frequent. These findings show that DRL can balance the benefit of preserving drug-sensitive cells to suppress resistance against the risk of unsafe tumour regrowth. Reward shaping further showed how treatment strategies could be adjusted to reflect different patient preferences. Together, these results provide a biologically informed approach for designing robust and patient-centred evolutionary therapies in fast-growing cancers such as NSCLC.
## Related Clinical Research

- [Translational modelling questions receptor‑occupancy‑based dosing for pembrolizumab (PD‑1 inhibito](https://medichelpline.com/clinical-feed/british-journal-of-cancer-0-translational-modelling-challenges-receptor-occupancy-based-dosing-of-pd-1.md)
- [Isoflavone-derived mitochondrial complex I inhibitor IV-16 shows potent activity against NSCLC](https://medichelpline.com/clinical-feed/pubmed-42217500.md) (DOI: 10.1016/j.bioorg.2026.110036)
- [Inhaled paclitaxel nanoagglomerate delivery for lung cancer: biodistribution and efficacy in biore](https://medichelpline.com/clinical-feed/pubmed-42476273.md) (DOI: 10.1016/j.ijpharm.2026.127213)
- [PD-L1 expression and survival in unresectable or recurrent gastric cancer treated with first-line](https://medichelpline.com/clinical-feed/british-journal-of-cancer-0-pd-l1-expression-and-survival-in-unresectable-recurrent-gastric-cancer-treated.md)
- [Oncolytic Virotherapy: Evolving from Cold-to-Hot to a Foundational Immuno‑Oncology Platform](https://medichelpline.com/clinical-feed/nature-reviews-clinical-oncology-0-beyond-cold-to-hot-oncolytic-virotherapy-as-the-next-cornerstone-of-immuno.md)

## Navigation
- [← Back to Oncology Feed](https://medichelpline.com/clinical-feed/oncology.md)
- [← All Clinical Specialties](https://medichelpline.com/clinical-feed.md)
## Medical & Regulatory Disclaimer

> [!CAUTION]
> MedicHelpline content is structured for research, educational, and professional discovery purposes. It does not constitute individual medical advice, clinical diagnosis, or treatment recommendations.
> Always verify dosing, contraindications, and regulatory alerts against official product labeling and primary regulatory sources before clinical decision-making.