<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
  <front>
    <journal-meta>
      <journal-id journal-id-type="nlm-ta">reapress</journal-id>
      <journal-id journal-id-type="publisher-id">null</journal-id>
      <journal-title>reapress</journal-title><issn pub-type="ppub">3042-2248</issn><issn pub-type="epub">3042-2248</issn><publisher>
      	<publisher-name>reapress</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">https://doi.org/10.48313/maa.v1i4.108</article-id>
      <article-categories>
        <subj-group subj-group-type="heading">
          <subject>Research Article</subject>
        </subj-group>
        <subj-group><subject>Dynamic treatment regime, Personalized medicine, Metaheuristic optimization, Reinforcement learning, Clinical decision support, Type 2 diabetes, Sepsis management.</subject></subj-group>
      </article-categories>
      <title-group>
        <article-title>Hybrid Metaheuristic–Reinforcement Learning Framework for Personalized Treatment Policy Optimization</article-title><subtitle>Hybrid Metaheuristic–Reinforcement Learning Framework for Personalized Treatment Policy Optimization</subtitle></title-group>
      <contrib-group><contrib contrib-type="author">
	<name name-style="western">
	<surname>Al-Qahtani</surname>
		<given-names>Reem </given-names>
	</name>
	<aff>Department of Computer Science, College of Computer and Information Sciences, King Saud University, Riyadh, Saudi Arabia.</aff>
	</contrib><contrib contrib-type="author">
	<name name-style="western">
	<surname>Nair</surname>
		<given-names>Arjun </given-names>
	</name>
	<aff>Department of Computer Science and Engineering, Indian Institute of Technology Delhi, New Delhi, India.</aff>
	</contrib></contrib-group>		
      <pub-date pub-type="ppub">
        <month>12</month>
        <year>2024</year>
      </pub-date>
      <pub-date pub-type="epub">
        <day>24</day>
        <month>12</month>
        <year>2024</year>
      </pub-date>
      <volume>1</volume>
      <issue>4</issue>
      <permissions>
        <copyright-statement>© 2024 reapress</copyright-statement>
        <copyright-year>2024</copyright-year>
        <license license-type="open-access" xlink:href="http://creativecommons.org/licenses/by/2.5/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</p></license>
      </permissions>
      <related-article related-article-type="companion" vol="2" page="e235" id="RA1" ext-link-type="pmc">
			<article-title>Hybrid Metaheuristic–Reinforcement Learning Framework for Personalized Treatment Policy Optimization</article-title>
      </related-article>
	  <abstract abstract-type="toc">
		<p>
			Personalized medicine increasingly demands dynamic, patient-specific treatment strategies that adapt to evolving clinical states, yet optimizing such strategies remains computationally intractable for conventional approaches. Dynamic Treatment Regimes (DTRs) formalize this challenge as sequential decision-making under uncertainty, but existing Reinforcement Learning (RL) methods for DTR optimization suffer from sample inefficiency, susceptibility to local optima, and safety concerns when trained on limited offline clinical data. This paper introduces Metaheuristic–Reinforcement learning (METRO-RL) optimizer, a novel two-stage hybrid framework that synergistically combines Differential Evolution (DE) for global policy search with a Deep Q-Network (DQN) variant for local policy refinement. In Stage 1, a modified DE algorithm with patient-state-aware mutation operators explores the policy parameter space globally, identifying promising policy regions while respecting clinical safety constraints. In Stage 2, a dueling DQN with Conservative Q-Learning (CQL) regularization fine-tunes the best candidate policies using prioritized experience replay from offline clinical datasets. A novel patient similarity kernel based on demographic features, clinical trajectory alignment via Dynamic Time Warping (DTW), and comorbidity profiles guides both stages, enabling effective knowledge transfer across patient subgroups. We evaluate METRO-RL on two clinically significant applications: Type 2 Diabetes (T2D) management using a simulated 10,000-patient cohort based on UVA/Padova simulator parameters and sepsis treatment in the Intensive Care Unit (ICU) using the MIMIC-IV database (approximately 18,500 sepsis episodes). On T2D management, METRO-RL achieves a composite glycemic control score of 0.782 ± 0.024, representing a 16.2% improvement over standard DQN and a 18.8% improvement over the clinician policy. For sepsis treatment, METRO-RL yields an estimated 90-day mortality of 18.7% ± 2.1%, an 11.4% relative reduction compared to the clinician policy. METRO-RL consistently outperforms standalone DQN, actor-critic, Batch-Constrained Q-Learning (BCQ), CQL, and evolutionary strategy baselines across all Off-Policy Evaluation (OPE) metrics while maintaining the lowest rate of clinically contraindicated actions (0.3%). 
		</p>
		</abstract>
    </article-meta>
  </front>
  <body></body>
  <back>
    <ack>
      <p>null</p>
    </ack>
  </back>
</article>