Reinforcement learning for chemotherapy scheduling in a stochastic tumor evolution model
Phys. Rev. E 113, 064404 – Published 5 June, 2026
DOI: https://doi.org/10.1103/cbkw-5s4b
Abstract
We present a -learning framework for optimizing chemotherapy dosing schedules in a stochastic finite-cell model of tumor evolution under drug-induced selection. The tumor consists of three competing subpopulations: a chemosensitive lineage, , and two single-drug-resistant lineages, and , each resistant to one of two cytotoxic agents, and . Drug administration is formulated as a discrete action space in a finite-state Markov process, where the state is defined by the composition constrained by a fixed population size . Tumor volume evolves according to a separate growth equation, with expansion rate proportional to the difference between the population-averaged fitness and a fixed microenvironmental baseline. Using -learning, we derive optimal dosing policies that balance therapeutic pressure with the evolutionary dynamics of resistance. The reward function is engineered to promote long-term coexistence among subpopulations, thereby delaying fixation of resistance by penalizing population imbalance. We analyze the structure of the optimal policies to (i) infer dominant evolutionary trajectories under treatment, (ii) quantify robustness to partial observability of both initial conditions and state updates, and (iii) construct simplified, symmetry-informed heuristics that approximate the full learned policy. Our results highlight the potential of model-free adaptive control strategies to steer tumor evolution away from drug resistance in the presence of biological stochasticity and information constraints.