The Performance of Different Artificial Intelligence Models in Predicting Breast Cancer among Individuals Having Type 2 Diabetes Mellitus

Meng-Hsuen Hsieh; Li-Min Sun; Cheng-Li Lin; Meng-Ju Hsieh; Chung Y Hsu; Chia-Hung Kao

doi:10.3390/cancers11111751

The Performance of Different Artificial Intelligence Models in Predicting Breast Cancer among Individuals Having Type 2 Diabetes Mellitus

Cancers (Basel). 2019 Nov 8;11(11):1751. doi: 10.3390/cancers11111751.

Authors

Meng-Hsuen Hsieh¹, Li-Min Sun^{2

3}, Cheng-Li Lin^{4

5}, Meng-Ju Hsieh⁶, Chung Y Hsu⁷, Chia-Hung Kao^{7

8

9}

Affiliations

¹ Department of Electrical Engineering and Computer Sciences, University of California, Berkeley, CA 94720, USA.
² Department of Radiation Oncology, Zuoying Branch of Kaohsiung Armed Forces General Hospital, Kaohsiung 813, Taiwan.
³ Institute of Medical Science and Technology, National Sun Yat-sen University, Kaohsiung 804, Taiwan.
⁴ Management Office for Health Data, China Medical University Hospital, Taichung 404, Taiwan.
⁵ College of Medicine, China Medical University, Taichung 404, Taiwan.
⁶ Department of Medicine, Poznan University of Medical Sciences, 60965 Poznan, Poland.
⁷ Graduate Institute of Biomedical Sciences, China Medical University, Taichung 404, Taiwan.
⁸ Department of Nuclear Medicine and PET Center, China Medical University Hospital, Taichung 404, Taiwan.
⁹ Department of Bioinformatics and Medical Engineering, Asia University, Taichung 404, Taiwan.

Abstract

Objective: Early reports indicate that individuals with type 2 diabetes mellitus (T2DM) may have a greater incidence of breast malignancy than patients without T2DM. The aim of this study was to investigate the effectiveness of three different models for predicting risk of breast cancer in patients with T2DM of different characteristics. Study design and methodology: From 2000 to 2012, data on 636,111 newly diagnosed female T2DM patients were available in the Taiwan's National Health Insurance Research Database. By applying their data, a risk prediction model of breast cancer in patients with T2DM was created. We also collected data on potential predictors of breast cancer so that adjustments for their effect could be made in the analysis. Synthetic Minority Oversampling Technology (SMOTE) was utilized to increase data for small population samples. Each datum was randomly assigned based on a ratio of about 39:1 into the training and test sets. Logistic Regression (LR), Artificial Neural Network (ANN) and Random Forest (RF) models were determined using recall, accuracy, F₁ score and area under the receiver operating characteristic curve (AUC). Results: The AUC of the LR (0.834), ANN (0.865), and RF (0.959) models were found. The largest AUC among the three models was seen in the RF model. Conclusions: Although the LR, ANN, and RF models all showed high accuracy predicting the risk of breast cancer in Taiwanese with T2DM, the RF model performed best.

Keywords: artificial neural network; breast cancer; logistic regression; random forest; type II diabetes mellitus.