Hybrid Metaheuristic–Reinforcement Learning Framework for Personalized Treatment Policy Optimization
Abstract
Personalized medicine increasingly demands dynamic, patient-specific treatment strategies that adapt to evolving clinical states, yet optimizing such strategies remains computationally intractable for conventional approaches. Dynamic Treatment Regimes (DTRs) formalize this challenge as sequential decision-making under uncertainty, but existing Reinforcement Learning (RL) methods for DTR optimization suffer from sample inefficiency, susceptibility to local optima, and safety concerns when trained on limited offline clinical data. This paper introduces Metaheuristic–Reinforcement learning (METRO-RL) optimizer, a novel two-stage hybrid framework that synergistically combines Differential Evolution (DE) for global policy search with a Deep Q-Network (DQN) variant for local policy refinement. In Stage 1, a modified DE algorithm with patient-state-aware mutation operators explores the policy parameter space globally, identifying promising policy regions while respecting clinical safety constraints. In Stage 2, a dueling DQN with Conservative Q-Learning (CQL) regularization fine-tunes the best candidate policies using prioritized experience replay from offline clinical datasets. A novel patient similarity kernel based on demographic features, clinical trajectory alignment via Dynamic Time Warping (DTW), and comorbidity profiles guides both stages, enabling effective knowledge transfer across patient subgroups. We evaluate METRO-RL on two clinically significant applications: Type 2 Diabetes (T2D) management using a simulated 10,000-patient cohort based on UVA/Padova simulator parameters and sepsis treatment in the Intensive Care Unit (ICU) using the MIMIC-IV database (approximately 18,500 sepsis episodes). On T2D management, METRO-RL achieves a composite glycemic control score of 0.782 ± 0.024, representing a 16.2% improvement over standard DQN and a 18.8% improvement over the clinician policy. For sepsis treatment, METRO-RL yields an estimated 90-day mortality of 18.7% ± 2.1%, an 11.4% relative reduction compared to the clinician policy. METRO-RL consistently outperforms standalone DQN, actor-critic, Batch-Constrained Q-Learning (BCQ), CQL, and evolutionary strategy baselines across all Off-Policy Evaluation (OPE) metrics while maintaining the lowest rate of clinically contraindicated actions (0.3%).
Keywords:
Dynamic treatment regime, Personalized medicine, Metaheuristic optimization, Reinforcement learning, Clinical decision support, Type 2 diabetes, Sepsis managementReferences
- [1] Collins, F. S., & Varmus, H. (2015). A new initiative on precision medicine. The new england journal of medicine, 372(9), 793. https://doi.org/10.1056/NEJMp1500523
- [2] Rajkomar, A., Dean, J., & Kohane, I. (2019). Machine learning in medicine. New england journal of medicine, 380(14), 1347–1358. https://doi.org/10.1056/NEJMra1814259
- [3] Murphy, S. A. (2003). Optimal dynamic treatment regimes. Journal of the royal statistical society series b: Statistical methodology, 65(2), 331–355. https://doi.org/10.1111/1467-9868.00389
- [4] Robins, J. M. (2004). Optimal structural nested models for optimal sequential decisions. In Proceedings of the second seattle symposium in biostatistics: Analysis of correlated data (pp. 189-326). New York, NY: Springer New York. https://doi.org/10.1007/978-1-4419-9076-1_11
- [5] Chakraborty, B., & Moodie, E. E. (2013). Statistical methods for dynamic treatment regimes: Reinforcement learning, causal inference, and personalized medicine. New York: Springer. https://doi.org/10.1007/978-1-4614-7428-9
- [6] Murphy, S. A. (2005). A generalization error for Q-learning. https://jmlr.org/papers/volume6/murphy05a/murphy05a.pdf
- [7] Watkins, C. J. C. H., & Dayan, P. (1992). Q-learning. Machine learning, 8(3), 279–292. https://doi.org/10.1007/BF00992698
- [8] Robins, J. M., Hernan, M. A., & Brumback, B. (2000). Marginal structural models and causal inference in epidemiology. Epidemiology, 11(5), 550–560. https://doi.org/10.1097/00001648-200009000-00011
- [9] Zhao, Y., Zeng, D., Rush, A. J., & Kosorok, M. R. (2012). Estimating individualized treatment rules using outcome weighted learning. Journal of the american statistical association, 107(499), 1106–1118. https://doi.org/10.1080/01621459.2012.695674
- [10] Raghu, A., Komorowski, M., Celi, L. A., Szolovits, P., & Ghassemi, M. (2017). Continuous state-space models for optimal sepsis treatment: A deep reinforcement learning approach. Machine learning for healthcare conference (pp. 147-163). PMLR. https://doi.org/10.48550/arXiv.1705.08422
- [11] Padmanabhan, R., Meskin, N., & Haddad, W. M. (2017). Reinforcement learning-based control of drug dosing for cancer chemotherapy treatment. Mathematical biosciences, 293, 11–20. https://doi.org/10.1016/j.mbs.2017.08.004
- [12] Shortreed, S. M., Laber, E., Lizotte, D. J., Stroup, T. S., Pineau, J., & Murphy, S. A. (2011). Informing sequential clinical decision-making through reinforcement learning: An empirical study. Machine learning, 84(1), 109–136. https://doi.org/10.1007/s10994-010-5229-0
- [13] Daskalaki, E., Diem, P., & Mougiakakou, S. G. (2013). An actor-critic based controller for glucose regulation in type 1 diabetes. Computer methods and programs in biomedicine, 109(2), 116–125. https://doi.org/10.1016/j.cmpb.2012.03.002
- [14] Gottesman, O., Johansson, F., Komorowski, M., Faisal, A., Sontag, D., Doshi-Velez, F., & Celi, L. A. (2019). Guidelines for reinforcement learning in healthcare. Nature medicine, 25(1), 16–18. https://doi.org/10.1038/s41591-018-0310-5
- [15] Levine, S., Kumar, A., Tucker, G., & Fu, J. (2020). Offline reinforcement learning: Tutorial, review, and perspectives on open problems. https://doi.org/10.48550/arXiv.2005.01643
- [16] Fujimoto, S., Meger, D., & Precup, D. (2019). Off-policy deep reinforcement learning without exploration. International conference on machine learning (pp. 2052-2062). PMLR. https://proceedings.mlr.press/v97/fujimoto19a/fujimoto19a.pdf
- [17] Kumar, A., Zhou, A., Tucker, G., & Levine, S. (2020). Conservative q-learning for offline reinforcement learning. Advances in neural information processing systems, 33, 1179–1191. https://dl.acm.org/doi/abs/10.5555/3495724.3495824
- [18] Achiam, J., Held, D., Tamar, A., & Abbeel, P. (2017). Constrained policy optimization. International conference on machine learning (pp. 22-31). PMLR. https://proceedings.mlr.press/v70/achiam17a.html
- [19] Storn, R., & Price, K. (1997). Differential evolution-A simple and efficient heuristic for global optimization over continuous spaces. Journal of global optimization, 11(4), 341–359. https://doi.org/10.1023/A:1008202821328
- [20] Das, S., & Suganthan, P. N. (2010). Differential evolution: A survey of the state-of-the-art. IEEE transactions on evolutionary computation, 15(1), 4–31. https://doi.org/10.1109/TEVC.2010.2059031
- [21] Petrovski, A., & McCall, J. (2001). Multi-objective optimisation of cancer chemotherapy using evolutionary algorithms. International conference on evolutionary multi-criterion optimization (pp. 531-545). Berlin, Heidelberg: Springer Berlin Heidelberg. https://doi.org/10.1007/3-540-44719-9_37
- [22] Holdsworth, C., Kim, M., Liao, J., & Phillips, M. H. (2010). A hierarchical evolutionary algorithm for multiobjective optimization in IMRT. Medical physics, 37(9), 4986–4997. https://doi.org/10.1118/1.3478276
- [23] Abo-Hammour, Z. S., Samhouri, A. D., & Mubarak, Y. (2014). Continuous genetic algorithm as a novel solver for Stokes and nonlinear Navier Stokes problems. Mathematical problems in engineering, 2014(1), 649630. https://doi.org/10.1155/2014/649630
- [24] Stanley, K. O., Clune, J., Lehman, J., & Miikkulainen, R. (2019). Designing neural networks through neuroevolution. Nature machine intelligence, 1(1), 24–35. https://doi.org/10.1038/s42256-018-0006-z
- [25] Jaderberg, M., Dalibard, V., Osindero, S., Czarnecki, W. M., Donahue, J., Razavi, A., ... & Kavukcuoglu, K. (2017). Population based training of neural networks. https://doi.org/10.48550/arXiv.1711.09846
- [26] Salimans, T., Ho, J., Chen, X., Sidor, S., & Sutskever, I. (2017). Evolution strategies as a scalable alternative to reinforcement learning. https://doi.org/10.48550/arXiv.1703.03864
- [27] Khadka, S., & Tumer, K. (2018). Evolution-guided policy gradient in reinforcement learning. Advances in neural information processing systems, 31. https://dl.acm.org/doi/10.5555/3326943.3327053
- [28] Robins, J. (1986). A new approach to causal inference in mortality studies with a sustained exposure period—application to control of the healthy worker survivor effect. Mathematical modelling, 7(9–12), 1393–1512. https://doi.org/10.1016/0270-0255(86)90088-6
- [29] Zhao, Y., Zeng, D., Socinski, M. A., & Kosorok, M. R. (2011). Reinforcement learning strategies for clinical trials in nonsmall cell lung cancer. Biometrics, 67(4), 1422–1433. https://doi.org/10.1111/j.1541-0420.2011.01572.x
- [30] Moodie, E. E. M., Dean, N., & Sun, Y. R. (2014). Q-learning: Flexible learning about useful utilities. Statistics in biosciences, 6(2), 223–243. https://doi.org/10.1007/s12561-013-9103-z
- [31] Lei, H., Nahum-Shani, I., Lynch, K., Oslin, D., & Murphy, S. A. (2012). A" SMART" design for building individualized treatment sequences. Annual review of clinical psychology, 8(1), 21–48. https://doi.org/10.1146/annurev-clinpsy-032511-143152
- [32] Almirall, D., Nahum-Shani, I., Sherwood, N. E., & Murphy, S. A. (2014). Introduction to SMART designs for the development of adaptive interventions: With application to weight loss research. Translational behavioral medicine, 4(3), 260–274. https://doi.org/10.1007/s13142-014-0265-0
- [33] Zhang, B., Tsiatis, A. A., Laber, E. B., & Davidian, M. (2013). Robust estimation of optimal dynamic treatment regimes for sequential treatment decisions. Biometrika, 100(3), 681–694. https://doi.org/10.1093/biomet/ast014
- [34] Zhao, Y.-Q., Zeng, D., Laber, E. B., & Kosorok, M. R. (2015). New statistical learning methods for estimating optimal dynamic treatment regimes. Journal of the american statistical association, 110(510), 583–598. https://doi.org/10.1080/01621459.2014.937488
- [35] Laber, E. B., & Zhao, Y.Q. (2015). Tree-based methods for individualized treatment regimes. Biometrika, 102(3), 501–514. https://doi.org/10.1093/biomet/asv028
- [36] Zhang, Y., Laber, E. B., Davidian, M., & Tsiatis, A. A. (2018). Interpretable dynamic treatment regimes. Journal of the american statistical association, 113(524), 1541–1549. https://doi.org/10.1080/01621459.2017.1345743
- [37] Tsiatis, A. A., Davidian, M., Holloway, S. T., & Laber, E. B. (2019). Dynamic treatment regimes: Statistical methods for precision medicine. CRC Press. https://doi.org/10.1201/9780429192692
- [38] Komorowski, M., Celi, L. A., Badawi, O., Gordon, A. C., & Faisal, A. A. (2018). The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care. Nature medicine, 24(11), 1716–1720. https://doi.org/10.1038/s41591-018-0213-5
- [39] Prasad, N., Cheng, L.F., Chivers, C., Draugelis, M., & Engelhardt, B. E. (2017). A reinforcement learning approach to weaning of mechanical ventilation in intensive care units. https://doi.org/10.48550/arXiv.1704.06300
- [40] Nemati, S., Ghassemi, M. M., & Clifford, G. D. (2016). Optimal medication dosing from suboptimal clinical examples: A deep reinforcement learning approach. 2016 38th annual international conference of the IEEE engineering in medicine and biology society (EMBC) (pp. 2978-2981). IEEE. https://doi.org/10.1109/EMBC.2016.7591355
- [41] Fox, I., Lee, J., Pop-Busui, R., & Wiens, J. (2020). Deep reinforcement learning for closed-loop blood glucose control. Machine learning for healthcare conference (pp. 508-536). PMLR. https://proceedings.mlr.press/v126/fox20a.html
- [42] Wu, Y., Tucker, G., & Nachum, O. (2019). Behavior regularized offline reinforcement learning. https://doi.org/10.48550/arXiv.1911.11361
- [43] Tessler, C., Mankowitz, D. J., & Mannor, S. (2018). Reward constrained policy optimization. https://doi.org/10.48550/arXiv.1805.11074
- [44] Huang, S., Papernot, N., Goodfellow, I., Duan, Y., & Abbeel, P. (2017). Adversarial attacks on neural network policies. https://doi.org/10.48550/arXiv.1702.02284
- [45] Hanset, A., Meskens, N., & Duvivier, D. (2010). Using constraint programming to schedule an operating theatre. 2010 IEEE workshop on health care management (WHCM) (pp. 1-6). IEEE. https://doi.org/10.1109/WHCM.2010.5441245
- [46] Anter, A. M., & Hassenian, A. E. (2019). CT liver tumor segmentation hybrid approach using neutrosophic sets, fast fuzzy c-means and adaptive watershed algorithm. Artificial intelligence in medicine, 97, 105–117. https://doi.org/10.1016/j.artmed.2018.11.007
- [47] Li, Y., Yao, J., & Yao, D. (2004). Automatic beam angle selection in IMRT planning using genetic algorithm. Physics in medicine & biology, 49(10), 1915–1932. https://doi.org/10.1088/0031-9155/49/10/007
- [48] Veng-Pedersen, P., Gobburu, J. V. S., Meyer, M. C., & Straughn, A. B. (2000). Carbamazepine level-A in vivo-in vitro correlation (IVIVC): A scaled convolution based predictive approach. Biopharmaceutics & drug disposition, 21(1), 1–6. https://doi.org/10.1002/1099-081X(200001)21:1%3C1::AID-BDD207%3E3.0.CO;2-D
- [49] Marchetti, G., Barolo, M., Jovanovic, L., Zisser, H., & Seborg, D. E. (2008). An improved PID switching control strategy for type 1 diabetes. IEEE transactions on biomedical engineering, 55(3), 857–865. https://doi.org/10.1109/TBME.2008.915665
- [50] Moles, C. G., Mendes, P., & Banga, J. R. (2003). Parameter estimation in biochemical pathways: A comparison of global optimization methods. Genome research, 13(11), 2467–2474. https://doi.org/10.1101/gr.1262503
- [51] Riquelme, N., Von Lücken, C., & Baran, B. (2015). Performance metrics in multi-objective optimization. 2015 Latin American computing conference (CLEI) (pp. 1-11). IEEE. https://doi.org/10.1109/CLEI.2015.7360024
- [52] Stanley, K. O., & Miikkulainen, R. (2002). Evolving neural networks through augmenting topologies. Evolutionary computation, 10(2), 99–127. https://doi.org/10.1162/106365602320169811
- [53] Stanley, K. O., D’Ambrosio, D. B., & Gauci, J. (2009). A hypercube-based encoding for evolving large-scale neural networks. Artificial life, 15(2), 185–212. https://doi.org/10.1162/artl.2009.15.2.15202
- [54] Pourchot, A., & Sigaud, O. (2018). CEM-RL: Combining evolutionary and gradient-based methods for policy search. https://doi.org/10.48550/arXiv.1810.01222
- [55] Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., & Meger, D. (2018). Deep reinforcement learning that matters. Proceedings of the AAAI conference on artificial intelligence (pp. 3207–3214). AAAI Press. https://doi.org/10.1609/aaai.v32i1.11694
- [56] Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., & Freitas, N. (2016). Dueling network architectures for deep reinforcement learning. International conference on machine learning (pp. 1995-2003). PMLR. https://proceedings.mlr.press/v48/wangf16.html
- [57] Schaul, T., Quan, J., Antonoglou, I., & Silver, D. (2015). Prioritized experience replay. https://doi.org/10.48550/arXiv.1511.05952
- [58] Precup, D., Sutton, R. S., & Singh, S. (2000). Eligibility traces for off-policy policy evaluation. Proceedings of the seventeenth international conference on machine learning (ICML) (pp. 759–766). Morgan Kaufmann. https://hdl.handle.net/20.500.14394/10401
- [59] Le, H., Voloshin, C., & Yue, Y. (2019). Batch policy learning under constraints. Proceedings of the 36th international conference on machine learning, 97, 3703–3712. https://proceedings.mlr.press/v97/le19a.html
- [60] Jiang, N., & Li, L. (2016). Doubly robust off-policy value evaluation for reinforcement learning. Proceedings of the 33rd international conference on machine learning (PP. 652–661 ). PMLR. https://proceedings.mlr.press/v48/jiang16.pdf
- [61] Kovatchev, B. P., Breton, M., Dalla Man, C., & Cobelli, C. (2009). In silico preclinical trials: A proof of concept in closed-loop control of type 1 diabetes. Journal of Diabetes science and technology, 3(1), 44–55. https://doi.org/10.1177/193229680900300106
- [62] Man, C. D., Micheletto, F., Lv, D., Breton, M., Kovatchev, B., & Cobelli, C. (2014). The UVA/PADOVA type 1 diabetes simulator: New features. Journal of Diabetes science and technology, 8(1), 26–34. https://doi.org/10.1177/1932296813514502
- [63] Holman, R. R., Paul, S. K., Bethel, M. A., Matthews, D. R., & Neil, H. A. W. (2008). 10-year follow-up of intensive glucose control in type 2 diabetes. New england journal of medicine, 359(15), 1577–1589. https://doi.org/10.1056/NEJMoa0806470
- [64] Johnson, A. E., Bulgarelli, L., Shen, L., Gayles, A., Shammout, A., Horng, S., ... & Mark, R. G. (2023). MIMIC-IV, a freely accessible electronic health record dataset. Scientific data, 10(1), 1. https://doi.org/10.1038/s41597-022-01899-x
- [65] Singer, M., Deutschman, C. S., Seymour, C. W., Shankar-Hari, M., Annane, D., Bauer, M., ... & Angus, D. C. (2016). The third international consensus definitions for sepsis and septic shock (Sepsis-3). JAMA, 315(8), 801-810. https://doi.org/10.1001/jama.2016.0287
- [66] Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., & Riedmiller, M. (2013). Playing Atari with deep reinforcement learning. https://doi.org/10.48550/arXiv.1312.5602
- [67] Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., & Kavukcuoglu, K. (2016). Asynchronous methods for deep reinforcement learning. Proceedings of the 33rd international conference on machine learning (pp. 1928–1937). PMLR. https://proceedings.mlr.press/v48/mniha16.html
- [68] Hansen, N. (2006). The CMA evolution strategy: A comparing review. Towards a new evolutionary computation: Advances in the estimation of distribution algorithms, 75–102. https://doi.org/10.1007/3-540-32494-1_4
- [69] Thomas, P., & Brunskill, E. (2016). Data-efficient off-policy policy evaluation for reinforcement learning. International conference on machine learning (pp. 2139-2148). PMLR. https://proceedings.mlr.press/v48/thomasa16.html
- [70] Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453. https://doi.org/10.1126/science.aax2342
- [71] Grote, T., & Berens, P. (2020). On the ethics of algorithmic decision-making in healthcare. Journal of medical ethics, 46(3), 205–211. https://doi.org/10.1136/medethics-2019-105586
- [72] Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., ... & Mordatch, I. (2021). Decision transformer: Reinforcement learning via sequence modeling. Advances in neural information processing systems, 34, 15084-15097. https://proceedings.neurips.cc/paper/2021/hash/7f489f642a0ddb10272b5c31057f0663-Abstract.html
- [73] Kostrikov, I., Nair, A., & Levine, S. (2021). Offline reinforcement learning with implicit q-learning. https://doi.org/10.48550/arXiv.2110.06169
- [74] Wang, Z., Hunt, J. J., & Zhou, M. (2022). Diffusion policies as an expressive policy class for offline reinforcement learning. https://doi.org/10.48550/arXiv.2208.06193