Reinforcement Learning for Personalised Antidepressant Sequencing in Treatment-Resistant Bipolar Depression: A Simulation-Based Policy Optimisation Study
Treatment-resistant bipolar depression (TRD-BD), defined in this simulation as failure to achieve remission after at least two adequate pharmacological trials for the current bipolar depressive episode, including at least one antidepressant trial delivered with mood-stabilising treatment, is managed through a sequence of uncertain decisions: continue, switch antidepressant class, or augment with a mood stabiliser or atypical antipsychotic. Current practice relies on consensus bipolar-disorder guidelines and sequential-care principles; the STAR*D study is referenced only as a structural example of multistep treatment sequencing in unipolar major depression, not as a validated bipolar-depression protocol. The sequencing problem is therefore a natural candidate for reinforcement learning (RL), but no RL framework has yet been developed specifically for personalised antidepressant sequencing in bipolar depression. We formulated antidepressant sequencing in TRD-BD as a finite-horizon Markov Decision Process and developed an offline RL framework for policy optimisation. A clinically grounded simulator encoded a 15-dimensional policy state comprising MADRS, YMRS, prior treatment failures, weeks on current treatment, side-effect burden, current drug class, a three-dimensional responder-profile covariate, age, and BD subtype, together with an 8-action treatment space (continue; switch to SSRI, SNRI, bupropion, or MAOI; augment with lithium, quetiapine, or lamotrigine). Although the responder profile represents an underlying biological trait, it was supplied to the policy as an oracle covariate in this proof-of-concept simulation; in real practice, it would usually be unobserved, making the clinical problem partially observable. The reward integrated MADRS improvement, remission, side-effect burden, switching cost, and manic-switch risk. A Deep Q-Network (DQN) was trained on 6172 transitions from 800 simulated trajectories generated by a stochastic behavioural clinician policy. Trajectories terminated at remission, manic switch, intolerable adverse effects, or the eight-stage horizon, yielding a mean logged trajectory length of 7.72 stages. Exact behavioural-policy action probabilities were recorded by the simulator and used for importance-weighted off-policy estimators. DQN was compared with Fitted Q-Iteration, the behavioural clinician policy, a fixed sequential protocol, and a random policy. The proposed DQN policy achieved an estimated discounted return of 41.6 (95% CI: 39.0 - 44.0), substantially exceeding Fitted Q-Iteration (24.6), the clinician policy (12.8), and the fixed protocol (12.1). The DQN policy achieved a remission rate of 61.8% (95% CI: 57.0% - 66.0%), nearly quadrupling the clinician policy’s 16.5% and reaching remission in fewer treatment stages (mean 6.5 vs 7.7). Mean MADRS reduction was 24.3 points versus 15.6 for the clinician policy. Bayesian posterior analysis assigned the DQN policy a 100.0% probability of being optimal. Policy analysis revealed that the learned strategy strongly favoured lithium augmentation selected in 88% of states over the clinician’s switch-heavy, lower-augmentation pattern. Within this synthetic environment, reinforcement learning derived a personalised sequencing policy that outperformed the simulated behavioural and fixed-protocol comparators. The learned preference for early lithium augmentation reflects the simulator parameters and should be interpreted as a proof-of-concept result rather than a clinical treatment recommendation. Validation with real TRD-BD trajectories and conservative offline-RL methods is required before any clinical inference or deployment.
Cite this paper
Filippis, R. D. and Foysal, A. A. (2026). Reinforcement Learning for Personalised Antidepressant Sequencing in Treatment-Resistant Bipolar Depression: A Simulation-Based Policy Optimisation Study. Open Access Library Journal, 13, e15677. doi: http://dx.doi.org/10.4236/oalib.1115677.
Tundo, A., Filippis, R.d. and Proietti, L. (2015) Pharmacologic Approaches to Treatment Resistant Depression: Evidences and Personal Experience. <i>World</i> <i>Journal</i> <i>of</i> <i>Psychiatry</i>, 5, 330-341. <br>https://doi.org/10.5498/wjp.v5.i3.330
Rush, A.J., Trivedi, M.H., Wisniewski, S.R., Nierenberg, A.A., Stewart, J.W., Warden, D., <i>et al</i>. (2006) Acute and Longer-Term Outcomes in Depressed Outpatients Requiring One or Several Treatment Steps: A STAR*D Report. <i>American Journal of Psychiatry</i>, 163, 1905-1917. <br>https://doi.org/10.1176/ajp.2006.163.11.1905
Hui Poon, S., Sim, K. and J. Baldessarini, R. (2015) Pharmacological Approaches for Treatment-Resistant Bipolar Disorder. <i>Current Neuropharmacology</i>, 13, 592-604. <br>https://doi.org/10.2174/1570159x13666150630171954
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A.A., Veness, J., Bellemare, M.G., <i>et al</i>. (2015) Human-Level Control through Deep Reinforcement Learning. <i>Nature</i>, 518, 529-533. <br>https://doi.org/10.1038/nature14236
Komorowski, M., Celi, L.A., Badawi, O., Gordon, A.C. and Faisal, A.A. (2018) The Artificial Intelligence Clinician Learns Optimal Treatment Strategies for Sepsis in Intensive Care. <i>Nature Medicine</i>, 24, 1716-1720. <br>https://doi.org/10.1038/s41591-018-0213-5
Prasad, N., Cheng, L.-F., Chivers, C., Draugelis, M. and Engelhardt, B.E. (2017) A Reinforcement Learning Approach to Weaning of Mechanical Ventilation in Intensive Care Units. <i>Proceedings of the </i>33<i>rd Conference on Uncertainty in Artificial Intelligence</i> (<i>UAI</i> 2017), Sydney, 11-15 August 2017.
Parbhoo, S., Bogojeska, J., Zazzi, M., Roth, V. and Doshi-Velez, F. (2017) Combining Kernel and Model Based Learning for HIV Therapy Selection. <i>AMIA Summits on Translational Science Proceedings</i>, San Francisco, 27-30 March 2017, 239-248.
Yatham, L.N., Kennedy, S.H., Parikh, S.V., <i>et al</i>. (2018) Canadian Network for Mood and Anxiety Treatments (CANMAT) and International Society for Bipolar Disorders (ISBD) 2018 Guidelines for the Management of Patients with Bipolar Disorder. <i>Bipolar Disorders</i>, 20, 97-170.
Levine, S., Kumar, A., Tucker, G. and Fu, J. (2020) Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems. <br>https://arxiv.org/abs/2005.01643
Jiang, N. and Li, L. (2016) Doubly Robust Off-Policy Value Evaluation for Reinforcement Learning. <i>Proceedings of the </i>33<i>rd International Conference on Machine Learning</i>, <i>PMLR</i>, Vol. 48, 652-661.
Kumar, A., Zhou, A., Tucker, G. and Levine, S. (2020) Conservative Q-Learning for Offline Reinforcement Learning. <i>Advances in Neural Information Processing Systems</i> 33, 6-12 December 2020, 1179-1191.
Pacchiarotti, I., Bond, D. J., Baldessarini, R. J., <i>et al</i>. (2013) The International Society for Bipolar Disorders (ISBD) Task Force Report on Antidepressant Use in Bipolar Disorders. <i>American Journal of Psychiatry</i>, 170, 1249-1262.
Thomas, P.S. and Brunskill, E. (2016) Data-Efficient Off-Policy Policy Evaluation for Reinforcement Learning. <i>Proceedings of the </i>33<i>rd International Conference on Machine Learning</i>, <i>PMLR</i>, Vol. 48, 2139-2148.
Raghu, A., Komorowski, M., Ahmed, I., Celi, L., Szolovits, P. and Ghassemi, M. (2017) Deep Reinforcement Learning for Sepsis Treatment. <br>https://arxiv.org/abs/1711.09602