Machine learning approach to differentiate iron deficiency anemia and thalassemia using random forest and gradient boosting algorithms

Machine Learning


Study population

This cross-sectional study was conducted from January 2015 to December 2019 at Songkranaga Lind Hospital, the largest tertiary hospital in southern Thailand. We evaluated initial visit data for 7488 patients with the following characteristics: (1) age >15 years, (2) Hb concentration <13 g/dL (men and menopausal women) or <12 g/dL (fertile women), (3) mean corpuscular volume (MCV) <80 fL, (4) available iron profile and ferritin level data, and (5) Hb and DNA analysis of Thal. Patients with anemia due to inflammation, transfusion-dependent Thal, pregnancy, or incomplete laboratory data were excluded. To rule out anemia due to inflammation and pregnancy, the hematologist reviewed the medical records to confirm the diagnosis of IDA and Thal and exclude patients with inflammation and infection.

Patients with serum ferritin level <30 ng/mL and transferrin saturation <16% were diagnosed with IDA15. All patients were diagnosed with Thal (TT and TI) using the following diagnostic criteria: Patients with Hb type A2A and Hb A2 levels ≥3.5% were diagnosed with β-TT. Hb type A2A, Hb A2 level < 3.5%、および DNA 分析後の α-Thal 変異陽性の患者は、α-TT と診断されました。 Hb タイプ EA および Hb E > Those with 10–35% were considered to have the Hb E trait. Patients diagnosed with TI showed Hb patterns such as A2FA, EFA, EE, A2AH, A2ABart'sH, CSA2AH, CSA2ABart'sH, EABart, EFABart, CSEABart, and CSEFABart. Furthermore, these patients had no history of blood transfusions and their symptoms were confirmed by DNA analysis. Abbreviation definitions and formal names are provided in the Appendix. Patients who met the criteria for IDA and Thal were diagnosed with IDA with Thal.

ethical approval

Ethical approval for this study was obtained from the Human Research Ethics Committee (HREC) of Prince of Songkla University Faculty of Medicine (REC 62-232-5-2). Because this study used anonymized data, HREC waived the requirement for informed consent.

experimental technology

Hematological characteristics were measured using an automatic hematology analyzer (XN3000, Sysmex Corporation, Kobe, Japan). Hb analysis was performed using capillary electrophoresis (CapillaryS2; Sebia, Lisses, France). Serum iron levels, total iron binding capacity, and ferritin levels were measured using an automatic analyzer (Cobas e411; Roche, Rotkreuz, Switzerland). DNA analysis of Thal was performed using polymerase chain reaction and reverse dot blot hybridization, as previously described.7,16.

statistical analysis

Baseline characteristics and hematologic features of TAR and IDA patients were compared using Pearson's chi-square test for categorical data and Kruskal-Wallis rank sum test for continuous data. P <0.05 was considered statistically significant. Complete blood count (CBC) data of patients diagnosed with Thal (TT and TI), IDA, and IDA with Thal (TT and TI) were divided into training and testing sets using a ratio of 80:20. Nine features were used in the machine learning method, including Hb level, hematocrit (Hct), MCV, mean corpuscular hemoglobin (MCH), mean corpuscular hemoglobin concentration (MCHC), red blood cell distribution width (RDW), RBC count, age, and gender.

Two types of models were built: a binary outcome model (Thal and IDA) and a multiclass outcome model (Thal, IDA, and IDA with Thal). Two ensemble machine learning classification methods were used to diagnose Thal and/or IDA. RF builds multiple decision trees on random subsets of the training dataset. This reduces the correlation between trees and avoids overfitting as each split selects a random subset of features. The final prediction is the majority vote from all trees. GB builds trees sequentially using information learned from previous trees. Theoretically, this improves accuracy compared to a single decision tree.17.

Our dataset was randomly split into training and testing datasets in a ratio of 80:20. To minimize overfitting and underfitting, model features in the training dataset were optimized using 10-fold cross-validation. We used the Latin hypercube method to sample 1000 sets of feature values ​​for each model. The best feature subset was the one that maximized the AUC-ROC on the validation data. The predictive model was then trained on the entire training data using the best features. Synthetic Minority Oversampling Technique (SMOTE) to address class imbalance18 was used to generate synthetic samples of the minority class, thereby effectively balancing the dataset before training the model.

The performance of the developed model was evaluated on the test dataset using various metrics such as accuracy, kappa coefficient, sensitivity, specificity, and AUC-ROC. Furthermore, we compared the diagnostic performance of GB and RF with that of prescriptions based on RBC indexes such as Hct/Hb, MCV/Hb, Keikhaei, Jayabose, Sirdah, Green and King, Mentzer, England, Fraser, Srivastava, Shine & Lal, Matos, Ricera, Kerman I, Kerman II, and Ehsani.8,19,20,21,22—By calculating the AUC-ROC difference using the method proposed by Delong et al.twenty three. All features and official definitions are provided in the appendix. The methodological flowchart of this study is shown in Supplementary Figure S1.

Analyzes were performed using R version 4.4.2.twenty four. The model specification and analysis process was performed in the following manner. neat model Version 1.2twenty five. The underlying analysis packages for RF and GB are: ranger Version 0.17.0 26 and xgboost Version 1.7.8.1 27respectively. Minority class oversampling was performed using: themeVersion 1.0.328. AUC-ROC calculations and comparisons were performed using pROC version 1.18.5.29.



Source link