Machine Learning for Analyzing Non-Countermeasure Factors Affecting Early Spread of COVID-19

Vito Janko; Gašper Slapničar; Erik Dovgan; Nina Reščič; Tine Kolenik; Martin Gjoreski; Maj Smerkol; Matjaž Gams; Mitja Luštrek

doi:10.3390/ijerph18136750

Machine Learning for Analyzing Non-Countermeasure Factors Affecting Early Spread of COVID-19

Int J Environ Res Public Health. 2021 Jun 23;18(13):6750. doi: 10.3390/ijerph18136750.

Authors

Vito Janko¹, Gašper Slapničar¹, Erik Dovgan¹, Nina Reščič¹, Tine Kolenik¹, Martin Gjoreski¹, Maj Smerkol¹, Matjaž Gams¹, Mitja Luštrek¹

Affiliation

¹ Jožef Stefan Institute, 1000 Ljubljana, Slovenia.

Abstract

The COVID-19 pandemic affected the whole world, but not all countries were impacted equally. This opens the question of what factors can explain the initial faster spread in some countries compared to others. Many such factors are overshadowed by the effect of the countermeasures, so we studied the early phases of the infection when countermeasures had not yet taken place. We collected the most diverse dataset of potentially relevant factors and infection metrics to date for this task. Using it, we show the importance of different factors and factor categories as determined by both statistical methods and machine learning (ML) feature selection (FS) approaches. Factors related to culture (e.g., individualism, openness), development, and travel proved the most important. A more thorough factor analysis was then made using a novel rule discovery algorithm. We also show how interconnected these factors are and caution against relying on ML analysis in isolation. Importantly, we explore potential pitfalls found in the methodology of similar work and demonstrate their impact on COVID-19 data analysis. Our best models using the decision tree classifier can predict the infection class with roughly 80% accuracy.

Keywords: COVID-19; feature correlation; feature significance; machine learning; risk factors.

Publication types

Research Support, Non-U.S. Gov't

MeSH terms

Algorithms
COVID-19*
Humans
Machine Learning
Pandemics
SARS-CoV-2

Grants and funding

P2-0209 (B)/Javna Agencija za Raziskovalno Dejavnost RS