Machine learning-based predictive model for brain syndrome after stroke: a risk stratification study

Machine Learning


Cohort Characteristics

Of the 571 screening records, 511 met the eligibility criteria. 178 (34.8%) developed CCS. Baseline comparisons between CCS and non-CCS groups showed significant differences in age (68.8 ± 7.9 vs 66.8 ± 8.2 years; p= 0.011), NIHSS (median 11 vs 9, p<0.001), D-Dimer (median 1.17 vs 0.79, p<0.001), CRP, HBA1C, and severe carotid stenosis (9.0% vs. 2.4%; p= 0.006) (Table 1).

Table 1 Baseline characteristics of study populations.

Model development and cross-validation

During the 5x cross-validation on the training set, SVM achieved the highest average Rocauc (0.799±0.056), but with a lower accuracy (0.662±0.004). Maximizing accuracy of random forest (0.819±0.018, AUC 0.794±0.042). XGBOOST and logistic regression balanced identification and accuracy (XGBOOST: AUC 0.779±0.042, accuracy 0.811±0.018; Logistic regression: AUC 0.785±0.052, accuracy 0.792±0.033). Deep neural network performance was moderate (AUC 0.759±0.041, accuracy 0.755±0.023). All AUC standard deviations are below 0.06, indicating stable discriminant ability for folding (Supplementary Table S1). Hyperparameter tuning did not enhance the test set AUC. The default model has been preserved. Supplementary Table S2 provides detailed explanations of the parameter settings and performance of each model before and after tuning.

Test Set Performance

Test cohort (n= 103), Xgboost achieves the highest identification (AUC 0.879; 95%CI 0.807–0.942), with accuracy 0.825, precision 0.844, recall 0.675, and F1 score 0.750. Random Forest followed (AUC 0.866, accuracy 0.845, precision 0.962, Recall 0.625, F1Score 0.758). SVM and logistic regression yielded an AUC of 0.853 and 0.818, respectively. Deep neural networks continued (AUC 0.817). Detailed indicators are summarized in Table 2 and the ROC curves shown in Figure 1 are shown.

Table 2 Comparison of model performance for five machine learning models.
Figure 1
Figure 1

ROC curves for five machine learning models. AUC: Area under the curve. ROC: Receiver operating characteristics. SVM: Support Vector Machine. xgboost: Extremely gradient boost.

The confusion matrix for each classifier further shows that XgBoost predicted 27 true positives, 58 true negatives, 5 false positives, 13 false negatives, and similar decompositions of random forests (25 TP, 62 TN, 1 FP, 15 FN), SVM (26 TP, 62 TN, 62 TN, 1 FP, 14 FN), and logistic regression (27 TP, 55 TN). Neural networks (29 TP, 52 TN, 11 FP, 11 FN) (Figure 2).

Figure 2
Figure 2

A confusion matrix of five machine learning models. Panel a – e is (a) Logistic regression, (b) Random Forest, (c) Supports vector machines (d)xgboost, and (e) Deep neural network. SVM: Support Vector Machine. xgboost: Extreme gradient boost; CCS: Brain-cardiac syndrome.

Calibration performance

Model calibrations were assessed using a Brier score across Hosmer-Lemeshow goodness-of-fit tests and all five machine learning algorithms. This analysis revealed significant differences in calibration quality between models (Table 3). The SVM demonstrated the best calibration performance at a Hosmer-Lemeshow P value of 0.246 (p>0.05, showing good calibration) and lowest briar score of 0.126. Random Forest also showed acceptable calibration (HL p-value = 0.153, Brier Score = 0.131).

Table 3 Model Calibration Performance: out of 5 Machine Learning Models.

In contrast, XGBoost showed that the calibration was insufficient despite achieving the highest identification performance with a Hosmer-Lemeshow P-Value <0.001 and a Brier score of 0.141. Similarly, deep neural networks showed severe misunderstanding (HL P value <0.001, Brier score = 0.178), and logistic regression showed moderate calibration problems (HL P value = 0.001, Brier score = 0.148).

Calibration plots (Figure 3) confirmed these findings visually. In SVM and Random Forest, we show that the predicted probability is closely matched with observed frequencies over decades, whereas Xgboost and deep neural networks showed systematic deviations from the diagonal lines of full calibration.

Figure 3
Figure 3

Calibration curves for five machine learning models. SVM: Support Vector Machine. xgboost: Extremely gradient boost.

Decision curve analysis

In decision curve analysis (Figure 4), all five models provide greater net benefits than the default strategy (“all treatment” or “no treatment”) over a wide range of clinically relevant threshold probability (0.1-0.8). Over the clinically relevant threshold range of 0.1-0.6, SVM and RF models deliver the highest net profits and consistently surpass both the “Treat-Oll” and “Treat-None” strategies. Xgboost offers slightly higher net profits only at very low thresholds (<0.15), reflecting the aggressive detection of true positives. Above 0.6, no model offered superiority over the default strategy.

Figure 4
Figure 4

Decision curve analysis of five machine learning models. SVM: Support Vector Machine. xgboost: Extremely gradient boost.

Threshold Optimization Analysis

The dataset showed moderate class imbalances (non-CCS:CCS = 1.96:1; CCS prevalence = 33.8% = 33.8%, 38.8% in test data), so we compared six threshold selection strategies. : 1 and 3:1.

The SVM model with an optimized threshold (0.576) provided the most balanced overall performance (accuracy = 87.4%, sensitivity = 70.0%, specificity = 98.4%, F1 score = 0.812) and outperformed all other models and strategy interventions. For high sensitivity applications, Xgboost adjusted with a 3:1 false negative penalty reached 95.0% sensitivity at a 0.019 threshold. Conversely, random forests retain the highest specificity (98.4%) at the default 0.50 threshold, making them suitable for check tests despite their low sensitivity (62.5%).

Throughout the model, the prevalence-based threshold consistently increased sensitivity (≈5-10% points) compared to the default, but did not maximize a single metric. Notably, Youden and F1 maximum thresholds converge on SVM, random forests, and Xgboost, indicating stable optimization despite modest imbalance. (Supplementary Table S3)

The importance of features

Following the importance analysis of features using SHAP, which identified D-Dimer as the most influential predictor, ACEI/ARB usage, HBA1C, CRP, and prothrombin time were closely followed. Additional important contributors include age, NIHSS score, and severe carotid stenosis. Figure 5a shows the importance plot of the SHAP feature, which ranks predictors based on average influence on model output. The SHAP summary plot (Figure 5B) further shows the distribution and orientation of these features, indicating high levels of D-Dimer, CRP, and HBA1C.

Figure 5
Figure 5

The importance of functionality and an overview of SHAP. (a) The importance of features ranked by average SHAP value (b)shap summary plot. Shap: Shapley Additive Description.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *