Improved disease diagnosis system for COVID-19 with data refactoring and handling methods

Ritesh Jha; Vandana Bhattacharjee; Abhijit Mustafi; Sudip Kumar Sahana

doi:10.3389/fpsyg.2022.951027

Improved disease diagnosis system for COVID-19 with data refactoring and handling methods

Front Psychol. 2022 Aug 12:13:951027. doi: 10.3389/fpsyg.2022.951027. eCollection 2022.

Authors

Ritesh Jha¹, Vandana Bhattacharjee¹, Abhijit Mustafi¹, Sudip Kumar Sahana¹

Affiliation

¹ Department of Computer Science and Engineering, Birla Institute of Technology, Mesra, Ranchi, India.

Abstract

The novel coronavirus illness (COVID-19) outbreak, which began in a seafood market in Wuhan, Hubei Province, China, in mid-December 2019, has spread to almost all countries, territories, and places throughout the world. And since the fault in diagnosis of a disease causes a psychological impact, this was very much visible in the spread of COVID-19. This research aims to address this issue by providing a better solution for diagnosis of the COVID-19 disease. The paper also addresses a very important issue of having less data for disease prediction models by elaborating on data handling techniques. Thus, special focus has been given on data processing and handling, with an aim to develop an improved machine learning model for diagnosis of COVID-19. Random Forest (RF), Decision tree (DT), K-Nearest Neighbor (KNN), Logistic Regression (LR), Support vector machine, and Deep Neural network (DNN) models are developed using the Hospital Israelita Albert Einstein (in São Paulo, Brazil) dataset to diagnose COVID-19. The dataset is pre-processed and distributed DT is applied to rank the features. Data augmentation has been applied to generate datasets for improving classification accuracy. The DNN model dominates overall techniques giving the highest accuracy of 96.99%, recall of 96.98%, and precision of 96.94%, which is better than or comparable to other research work. All the algorithms are implemented in a distributed environment on the Spark platform.

Keywords: COVID-19; classification; data augmentation; data pre-processing; disease diagnosis.