Using a dataset containing chest computed tomography severity score data to compare machine learning algorithms for predicting COVID-19 mortality

Machine Learning


A total of 6,854 suspected cases were referred to Ayatollah Talegani Hospital, and 815 RT-PCR-positive patients remained on record after applying exclusion criteria. Overall, 54.85% of his enrolled patients were male, and he had a mean age of 57.22 ± 16.76 years in the study population. As mentioned earlier, the deceased group contained only his 108 records (13%) and his SMOTE method was used to balance these data. After rebalancing the dataset, the number of records in this class increased to 707.

Feature selection

Twenty-seven features were selected as the most important and relevant predictors using the chi-square independence test. These features include demographics, risk factors, clinical manifestations, laboratory results, and CT-SS data. A list of the most important variables and the results of the chi-square independence test are shown in Table 3. The table also shows the mean reduction rate of gini and importance scores for these variables, calculated using the XGBoost test and the random forest test. Listed. Descriptive statistics for these features are shown in Table 4. In this study, the most relevant features included age, consistent with other studies reporting several important clinical predictors of mortality in patients with COVID-19.6,11,15,16,17,26,27,28,29,30sex6,16,17,26,29,30,31,32,33,34dry cough6,11,14,17,27,29,32,33,35 Clinical manifestations include underlying diseases such as cardiovascular disease6,15,17,27,28,34,36,37high blood pressure6,15,17,27,29,30,34,36Diabetes6,15,16,17neurological disorders6,16,17cancer6,17,26,29,37serum creatinine and other laboratory indicators6,17red blood cells6WBC6,29,35hematocrit6absolute lymphocyte count6,14,17,27,31,33,34the absolute number of neutrophils6,14,15,17,27,28,33,35,36,calcium6, 11, 33phosphor6blood urea nitrogen6,14,33total bilirubin6,35serum albumin6,14,29,33,34,glucose6,17creatinine kinase6,11,14,29,34,35activated partial platelet formation time6prothrombin time6,34hypersensitive troponin6,17,28and CT-SS as imaging manifestations6,28. These predictors were used as inputs to develop an ML-based model for mortality prediction of COVID-19 patients.

Table 3 Importance scores, mean percentage reductions in Gini values, and statistical significance levels of the most important variables for predicting COVID-19 mortality calculated using XGBoost, random forest, and chi-square tests.
Table 4 Descriptive statistics of the most important variables for predicting mortality in COVID-19 patients.

On the other hand, smoking6,15,17,28,30alcohol/addiction6, 17, 30sore throat6,15,16,17,26,27,31,33,38muscle pain and fatigue6,14,15,16,17,26,28,34diarrhea, gastrointestinal symptoms6,14,16,17,29,30,36headache6,11,17,26,30,31,37platelet count6, 14, 28, 29and alanine aminotransferase (ALT)6, 14, 29, 31 These features were irrelevant for predicting COVID-19 mortality. Despite the clinical importance of these parameters for treatment success and mortality prediction, many of them can be excluded from the ML analysis, allowing fewer elements to perform mortality prediction with the same accuracy. There is a nature.

Evaluation of the development model

In this study, a COVID-19 mortality prediction model was developed using eight ML algorithms including J48, SVM, MLP, k-NN, NB, LR, RF and XGBoost. These predictive models were built using the best feature subsets determined in the previous step. ML algorithms were trained using the same dataset. The performance of these models was evaluated using sensitivity, specificity, accuracy, precision, and AUC metrics. Table 5 shows the performance evaluation results of the developed model.

Table 5 Performance of ML algorithms for mortality prediction of COVID-19 patients.

The results showed that the RF algorithm outperformed other ML algorithms in predicting mortality in COVID-19 patients. The sensitivity, specificity, precision, precision, F1 score, and AUC of the RF algorithm were 100.0%, 94.5%, 97.2%, 94.8%, 97.3%, and 99.9%, respectively. Figure 3 shows a comparison of the areas under the ROC curves of the developed ML algorithms.

Figure 3
Figure 3

ROC curve of ML algorithm.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *