Antidepressant safety surveillance in bipolar disorder requires large, multi-site patient populations to achieve reliable pharmacovigilance signal detection yet the raw patient data necessary to train centralized machine learning models are subject to strict privacy regulations (GDPR, HIPAA) that prohibit cross-site data sharing in most jurisdictions. Federated learning (FL) offers a principled solution: locally trained models share only gradient updates or model weights with a central aggregation server, enabling collaborative learning across institutions without any raw data leaving the originating site. No federated learning framework has been applied to antidepressant safety surveillance in BD, and the privacy-utility tradeoff of differential privacy augmentation in this clinical context remains uncharacterised. We implemented and evaluated a Federated Averaging (FedAvg) framework across five fully simulated, heterogeneous bipolar-disorder registry sites (total N = 800; Sites A-E represented academic, community, European, primary-care, and specialised bipolar-clinic settings). A multilayer perceptron was trained locally for 15 communication rounds with five local epochs per round. Comparators included local-only models, centralised MLP and XGBoost reference models, FedProx, and differentially private FedAvg (DP-FedAvg). A privacy-utility analysis examined seven nominal privacy settings. The composite outcome was a simulated SSRI-associated safety event within six months of initiation, comprising mood switch, cycle acceleration, or unplanned discontinuation. In the held-out simulated test set, FedAvg achieved AUC = 0.957 (95% CI: 0.920 - 0.987), F1 = 0.842 (CI: 0.750 - 0.925), sensitivity = 0.814, and specificity = 0.936. Its point-estimate AUC was slightly higher than the centralised MLP (0.931) and similar to centralised XGBoost (0.954), FedProx (0.956), and local-only averaging (0.950). These small differences should be interpreted as performance equivalence within uncertainty, not evidence that federation intrinsically outperforms centralised training. FedAvg converged within approximately 12 communication rounds, and the simulated privacy-noise analysis showed limited AUC degradation over the tested range. Within this proof-of-concept simulation, federated training preserved predictive performance without sharing raw records. The findings support evaluation on real multi-site bipolar-disorder registries but do not by themselves establish clinical validity, formal GDPR/HIPAA compliance, or deployability.
Cite this paper
Filippis, R. D. and Foysal, A. A. (2026). Federated Learning for Privacy-Preserving Antidepressant Safety Surveillance across Multi-Site Bipolar Disorder Registries. Open Access Library Journal, 13, e15675. doi: http://dx.doi.org/10.4236/oalib.1115676.
Pacchiarotti, I., Bond, D.J., Baldessarini, R.J., Nolen, W.A., Grunze, H., Licht, R.W. and Vieta, E. (2013) The ISBD Task Force Report on Antidepressant Use in Bipolar Disorders. <i>American Journal of Psychiatry</i>, 170, 1249-1262.
(2016) Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 on the Protection of Natural Persons with Regard to the Processing of Personal Data (General Data Protection Regulation).
Wu, X.Y., Yao, X. and Wang, C.L. (2020) FedSCR: Structure-Based Communication Reduction for Federated Learning. <i>IEEE Transactions on Parallel and Distributed Systems</i>, 32, 1565-1577.
Rieke, N., Hancox, J., Li, W., Milletarì, F., Roth, H.R., Albarqouni, S., <i>et al</i>. (2020) The Future of Digital Health with Federated Learning. <i>npj Digital Medicine</i>, 3, Article No. 119. <br>https://doi.org/10.1038/s41746-020-00323-1
Li, T., Sahu, A.K., Zaheer, M., Sanjabi, M., Talwalkar, A. and Smithy, V. (2019). Feddane: A Federated Newton-Type Method. 2019 53<i>rd Asilomar Conference on Signals</i>, <i>Systems</i>, <i>and Compu</i><i>ters</i>, Pacific Grove, 3-6 November 2019. <br>https://doi.org/10.1109/ieeeconf44664.2019.9049023
Dwork, C., McSherry, F., Nissim, K. and Smith, A. (2006) Calibrating Noise to Sensitivity in Private Data Analysis. In: Halevi, S. and Rabin, T., Eds., <i>Lecture Notes in Computer Science</i>, Springer, 265-284. <br>https://doi.org/10.1007/11681878_14
Abadi, M., Chu, A., Goodfellow, I., McMahan, H.B., Mironov, I., Talwar, K., <i>et al</i>. (2016) Deep Learning with Differential Privacy. <i>Proceedings of the </i>2016<i> ACM SIGSAC Conference on Computer and Communications Security</i>, Vienna, 24-28 October 2016, 308-318. <br>https://doi.org/10.1145/2976749.2978318
Brisimi, T.S., Chen, R., Mela, T., Olshevsky, A., Paschalidis, I.C. and Shi, W. (2018) Federated Learning of Predictive Models from Federated Electronic Health Records. <i>International Journal of Medical Informatics</i>, 112, 59-67. <br>https://doi.org/10.1016/j.ijmedinf.2018.01.007
Pfohl, S.R., Foryciarz, A. and Shah, N.H. (2021) An Empirical Characterization of Fair Machine Learning for Clinical Risk Prediction. <i>Journal of Biomedical Informatics</i>, 113, Article 103621. <br>https://doi.org/10.1016/j.jbi.2020.103621
Dubey, P., Dubey, P. and Bokoro, P.N. (2025) Federated Learning for Privacy-Enhanced Mental Health Prediction with Multimodal Data Integration. <i>Computer Met</i><i>hods in Biomechanics and Biomedical Engineering: Imaging & Visualization</i>, 13, Article 2509672. <br>https://doi.org/10.1080/21681163.2025.2509672
Yatham, L.N., Kennedy, S.H., Parikh, S.V., Schaffer, A. and McIntyre, R. (2018) CANMAT and ISBD 2018 Guidelines for the Management of Patients with Bipolar Disorder. <i>Bipolar Disorders</i>, 20, 97-170.
Kass, R.E. and Raftery, A.E. (1995) Bayes Factors. <i>Journal of the American Statistical Association</i>, 90, 773-795. <br>https://doi.org/10.1080/01621459.1995.10476572
Chen, T. and Guestrin, C. (2016) XGBoost: A Scalable Tree Boosting System. <i>Proceedings of the</i> 22<i>nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</i>, San Francisco, 13-17 August 2016, 785-794. <br>https://doi.org/10.1145/2939672.2939785
Lundberg, S.M. and Lee, S.-I. (2017) A Unified Approach to Interpreting Model Predictions. <i>Advances in Neural Information Processing Systems</i>, 30, 4765-4774.
DeLong, E.R., DeLong, D.M. and Clarke-Pearson, D.L. (1988) Comparing the Areas under Two or More Correlated Receiver Operating Characteristic Curves: A Nonparametric Approach. <i>Biometrics</i>, 44, 837-844. <br>https://doi.org/10.2307/2531595
He, K., Zhang, X., Ren, S. and Sun, J. (2016) Deep Residual Learning for Image Recognition. 2016<i> IEEE Conference on Computer Vision and Pattern Recognition</i> (<i>CVPR</i>), Las Vegas, 27-30 June 2016, 770-778. <br>https://doi.org/10.1109/cvpr.2016.90