An Explainable Machine Learning Pipeline for Stroke Prediction on Imbalanced Data

Christos Kokkotis; Georgios Giarmatzis; Erasmia Giannakou; Serafeim Moustakidis; Themistoklis Tsatalas; Dimitrios Tsiptsios; Konstantinos Vadikolias; Nikolaos Aggelousis

doi:10.3390/diagnostics12102392

An Explainable Machine Learning Pipeline for Stroke Prediction on Imbalanced Data

Diagnostics (Basel). 2022 Oct 1;12(10):2392. doi: 10.3390/diagnostics12102392.

Authors

Christos Kokkotis¹, Georgios Giarmatzis¹, Erasmia Giannakou¹, Serafeim Moustakidis², Themistoklis Tsatalas³, Dimitrios Tsiptsios⁴, Konstantinos Vadikolias⁴, Nikolaos Aggelousis¹

Affiliations

¹ Department of Physical Education and Sport Science, Democritus University of Thrace, 69100 Komotini, Greece.
² AIDEAS OÜ, Narva mnt 5, 10117 Tallinn, Estonia.
³ Department of Physical Education and Sport Science, University of Thessaly, 38221 Trikala, Greece.
⁴ Department of Neurology, School of Medicine, University Hospital of Alexandroupolis, Democritus University of Thrace, 68100 Alexandroupolis, Greece.

Abstract

Stroke is an acute neurological dysfunction attributed to a focal injury of the central nervous system due to reduced blood flow to the brain. Nowadays, stroke is a global threat associated with premature death and huge economic consequences. Hence, there is an urgency to model the effect of several risk factors on stroke occurrence, and artificial intelligence (AI) seems to be the appropriate tool. In the present study, we aimed to (i) develop reliable machine learning (ML) prediction models for stroke disease; (ii) cope with a typical severe class imbalance problem, which is posed due to the stroke patients' class being significantly smaller than the healthy class; and (iii) interpret the model output for understanding the decision-making mechanism. The effectiveness of the proposed ML approach was investigated in a comparative analysis with six well-known classifiers with respect to metrics that are related to both generalization capability and prediction accuracy. The best overall false-negative rate was achieved by the Multi-Layer Perceptron (MLP) classifier (18.60%). Shapley Additive Explanations (SHAP) were employed to investigate the impact of the risk factors on the prediction output. The proposed AI method could lead to the creation of advanced and effective risk stratification strategies for each stroke patient, which would allow for timely diagnosis and the right treatments.

Keywords: clinical data; interpretation; machine learning; prognosis; stroke.

Grants and funding

MIS 5047286/Greek and European funds (EYD-EPANEK)