Replicability and reproducibility of predictive models for diagnosis of depression among young adults using Electronic Health Records

David Nickson; Henrik Singmann; Caroline Meyer; Carla Toro; Lukasz Walasek

doi:10.1186/s41512-023-00160-2

Replicability and reproducibility of predictive models for diagnosis of depression among young adults using Electronic Health Records

Diagn Progn Res. 2023 Dec 5;7(1):25. doi: 10.1186/s41512-023-00160-2.

Authors

David Nickson¹, Henrik Singmann², Caroline Meyer³, Carla Toro³, Lukasz Walasek⁴

Affiliations

¹ WMG, University of Warwick, Coventry, UK. david.nickson@warwick.ac.uk.
² Department of Experimental Psychology, University College London, London, UK.
³ Warwick Medical School, University of Warwick, Coventry, UK.
⁴ Department of Psychology, University of Warwick, Coventry, UK.

Abstract

Background: Recent advances in machine learning combined with the growing availability of digitized health records offer new opportunities for improving early diagnosis of depression. An emerging body of research shows that Electronic Health Records can be used to accurately predict cases of depression on the basis of individual's primary care records. The successes of these studies are undeniable, but there is a growing concern that their results may not be replicable, which could cast doubt on their clinical usefulness.

Methods: To address this issue in the present paper, we set out to reproduce and replicate the work by Nichols et al. (2018), who trained predictive models of depression among young adults using Electronic Healthcare Records. Our contribution consists of three parts. First, we attempt to replicate the methodology used by the original authors, acquiring a more up-to-date set of primary health care records to the same specification and reproducing their data processing and analysis. Second, we test models presented in the original paper on our own data, thus providing out-of-sample prediction of the predictive models. Third, we extend past work by considering several novel machine-learning approaches in an attempt to improve the predictive accuracy achieved in the original work.

Results: In summary, our results demonstrate that the work of Nichols et al. is largely reproducible and replicable. This was the case both for the replication of the original model and the out-of-sample replication applying NRCBM coefficients to our new EHRs data. Although alternative predictive models did not improve model performance over standard logistic regression, our results indicate that stepwise variable selection is not stable even in the case of large data sets.

Conclusion: We discuss the challenges associated with the research on mental health and Electronic Health Records, including the need to produce interpretable and robust models. We demonstrated some potential issues associated with the reliance on EHRs, including changes in the regulations and guidelines (such as the QOF guidelines in the UK) and reliance on visits to GP as a predictor of specific disorders.

Keywords: Depression; Electronic health records; Machine learning; Predictive modelling; Replicability; Reproducibility.

Grants and funding

2300953/Engineering and Physical Sciences Research Council