A study on the effectiveness of machine learning models for hepatitis prediction

Machine Learning


Table 2 provides a comprehensive breakdown of the various variables included in this study, which encompass both categorical and continuous data. For categorical variables, Table 2 shows the frequency and percentage of individuals within each category. For instance, in terms of sex, there are 139 males (89.7%) and 16 females (10.3%) in the study. Regarding steroid use, 76 individuals (49.0%) reported using steroids, while 79 individuals (51.0%) did not. When it comes to fatigue, 130 individuals (83.9%) experienced fatigue, whereas 25 individuals (16.1%) did not. In terms of liver enlargement, 24 individuals (15.5%) had an enlarged liver, while 131 individuals (84.5%) did not. Similarly, 51 individuals (32.9%) had ascites, and 104 individuals (67.1%) did not. The histology variable shows that 85 individuals (54.8%) had histological findings, while 70 individuals (45.2%) did not. The table also includes variables like spiders, bilirubin, alkaline phosphate, aspartate transaminase, albumin, and pro-time, though no specific frequencies or percentages are provided for these variables. The continuous variables, such as age, bilirubin, alkaline phosphate, aspartate transaminase, albumin, and pro-time, are listed with their respective ranges, but again, no specific frequencies or percentages are provided. This detailed table helps to organize and summarize the key characteristics of the study population, offering valuable insights into the distribution of various medical factors.

Table 2 Demographic characteristics of independent variables.

Table 3 breaks down various categorical variables such as sex, steroid use, spleen palpability, ascites, varices, histology, and liver enlargement, illustrating their connection to the outcomes of death (D, L = 1) and survival (D, L = 2). Each variable is accompanied by the number of individuals who died and survived, as well as the total count for each category. For example, among the male participants, 32 individuals died and 107 survived, whereas all 16 female participants survived, with no deaths recorded. In terms of steroid use, 20 individuals who used steroids died, while 56 survived, compared to the 12 deaths and 67 survivors among those who did not use steroids. The Table 3 also highlights the impact of other conditions such as spleen palpability, ascites, varices, histology, and liver enlargement on survival outcomes. For instance, among individuals with palpable spleens, 12 died and 18 survived, while those with non-palpable spleens experienced 20 deaths and 105 survivors. Additionally, no females in the study died, a noteworthy observation in understanding gender differences in hepatitis outcomes. The data also shows that a majority of individuals did not have liver enlargement, with 131 individuals without the condition compared to just 24 with it. Furthermore, a larger portion of the population did not use steroids, indicating a potential factor influencing survival. This comprehensive breakdown enables a deeper exploration of the correlation between medical conditions and survival rates, contributing to a better understanding of the factors affecting hepatitis outcomes.

Table 3 Crosstabulation of Independent Variables with Hepatitis Outcome (D, L).

Table 4 presents the results of the Boruta algorithm, highlighting the importance of various attributes in relation to the studied outcomes. It offers an integrated assessment of feature relevance for Hepatitis by combining data-driven insights from the algorithm with clinical domain expertise. The analysis highlights several variables that are crucial in understanding the progression and severity of hepatitis, especially in cases that advance to liver cirrhosis. Among the most important features are Ascites, Varices, Bilirubin, Age, Spiders, and Alkaline phosphatase, all of which received high importance scores (above 0.85). These features are well-established markers of liver dysfunction. For instance, the presence of ascites and varices indicates advanced liver damage and portal hypertension—key complications in decompensated liver disease. Elevated bilirubin levels and prolonged prothrombin time reflect impaired liver function, while clinical signs such as spider angiomas result from hormonal imbalances due to hepatic insufficiency. Age also plays a significant role, as older patients are at greater risk of fibrosis progression and hepatocellular carcinoma. Several features were categorized as tentative, showing moderate importance. These include Prothrombin time, Antiviral treatment, Liver Firmness, and Albumin. Although these variables are clinically relevant, their statistical contribution in the model was less consistent, possibly due to redundancy with stronger predictors or variability in clinical measurement. For example, albumin is a useful marker of liver synthetic function but may be influenced by multiple external factors, reducing its predictive clarity in the dataset. Similarly, antiviral use reflects treatment history, which can vary independently of disease severity. On the other hand, features such as Aspartate (AST), Liver Big, Fatigue, and Histology were considered unimportant due to their lower relevance in both the clinical and algorithmic contexts. While AST is a liver enzyme often elevated in hepatitis, its predictive value may be limited due to overlap with other liver markers. Fatigue is a common but nonspecific symptom, and liver enlargement is not unique to hepatitis-related conditions. Histology data, if inconsistently recorded, may not provide reliable input for model-based predictions.

Table 4 Overview of the Boruta algorithm results.

Overall, the table summarizes the Boruta algorithm’s findings, revealing which variables are most influential in the model and which may not be as valuable for further analysis, while Fig. 2 visually represents these important features, highlighting the key variables that significantly contribute to the analysis and model prediction. Figure 3 illustrates Random Forest’s feature importance, reinforcing these findings. Together, they provide a comprehensive view of the variables deemed crucial by the Boruta algorithm, guiding further exploration and model refinement.

Fig. 2
figure 2

Feature selection using Boruta algorithm for Hepatitis.

Fig. 3
figure 3

Random forest feature importance plot.

Performance evaluation of considered models

Table 5 presents the performance metrics for various classifiers, showcasing their accuracy, precision, sensitivity, specificity, and F1 score, along with their corresponding 95% confidence intervals (CIs).

Table 5 Comparison of various machine learning methods for hepatitis prediction.

In the evaluation of machine learning models for hepatitis prediction, Logistic Regression (LR) remains a strong baseline due to its simplicity and transparency. With an accuracy of 85.00% (95% CI 78.78%–91.22%), precision of 94.03% (CI 88.66%–99.39%), and F1 score of 91.30% (CI 85.78%–96.82%), LR delivers consistently reliable performance. Its sensitivity is notably high at 88.73% (CI 80.71%–96.76%), though specificity is modest at 55.56% (CI 38.30%–72.82%), indicating room for improvement in identifying negative cases. Nevertheless, its interpretability makes LR particularly valuable in clinical settings where model transparency is essential. Support Vector Machine (SVM) offers strong detection capabilities, especially for positive cases. It achieves a sensitivity of 89.71% (CI 82.13%–97.29%) and precision of 91.04% (CI 84.86%–97.22%), resulting in an F1 score of 90.38% (CI 84.42%–96.35%). However, its specificity is limited to 50.00% (CI 32.66%–67.34%), suggesting a higher rate of false positives. This makes SVM suitable in contexts where maximizing the identification of actual hepatitis cases outweighs the cost of occasional false alarms.

Random Forest (RF) emerges as the best-performing model across most metrics. It achieves the highest accuracy at 92.42% (CI 88.25%–96.59%), precision at 96.77% (CI 93.99%–99.55%), sensitivity at 95.24% (CI 89.92%–100.00%), and F1 score at 96.00% (CI 92.54%–99.46%). These narrow confidence intervals highlight RF’s consistency and robustness. The model’s only drawback is its low specificity of 33.33% (CI 15.79%–50.87%), suggesting a tendency to misclassify negatives as positives. Despite this, RF’s overall performance makes it the most comprehensive and dependable model for hepatitis detection when recall and precision are prioritized.

K-Nearest Neighbors (KNN) stands out for its high specificity of 87.76% (CI 78.79%–96.73%), making it effective in ruling out non-hepatitis cases. However, its lower precision (71.43%, CI 60.24%–82.62%) and sensitivity (78.95%, CI 67.50%–90.40%) result in a modest F1 score of 75.00% (CI 64.12%–85.88%). KNN is more suited to tasks where minimizing false positives is more important than detecting all true positives. Artificial Neural Networks (ANN) provide a well-rounded performance, with precision and sensitivity both at 86.96% (CIs: 79.03%–94.89% and 76.67%–97.25%, respectively), and an F1 score of 86.96% (CI 78.40%–95.52%). However, accuracy (80.65%, CI 73.92%–87.38%) and specificity (62.50%, CI 45.73%–79.27%) are moderate. This makes ANN valuable for modeling complex patterns, though less optimal when higher accuracy or specificity is critical.

AdaBoost excels in specificity with the highest score of 95.65% (CI 89.45%–100.00%), making it a top choice when reducing false positives is a priority. However, with a sensitivity of 50.00% (CI 32.69%–67.31%), precision of 76.92% (CI 61.68%–92.16%), and an F1 score of 61.54% (CI 45.08%–78.00%), it is less effective at identifying actual hepatitis cases, limiting its applicability in situations where missed diagnoses are unacceptable. XGBoost offers a balanced performance across most metrics, with an accuracy of 85.07% (CI 78.90%–91.23%), sensitivity of 77.78% (CI 65.42%–90.14%), specificity of 87.76% (CI 78.79%–96.73%), and an F1 score of 73.68% (CI 62.32%–85.04%). Its precision is relatively lower at 70.00% (CI 58.06%–81.94%), but its balanced nature and stability make it a viable option when both sensitivity and specificity need to be reasonably maintained.

The Fig. 4 compares seven ML models across five metrics. Random Forest shows the highest accuracy, precision, sensitivity, and F1-score, but low specificity. AdaBoost has high specificity but low sensitivity. Logistic Regression and SVM offer balanced performance. KNN, ANN, and XGBoost perform moderately, each with trade-offs across metrics.

Fig. 4
figure 4

Comparison of ML models performance.

The inclusion of 95% confidence intervals offers essential insight into the stability and reliability of each model’s performance. Narrow intervals, such as those seen with RF and SVM, suggest high consistency, while wider intervals, as observed with AdaBoost and ANN, indicate variability and potential uncertainty in performance. RF clearly leads in accuracy, sensitivity, precision, and F1 score, affirming its superiority for hepatitis prediction, especially in tasks demanding high recall and precise classification of positives. However, its low specificity warrants caution where false positives can have significant consequences. SVM and LR present themselves as strong alternatives, balancing high performance with more interpretability and slightly better specificity. ANN provides moderate versatility, while KNN and XGBoost cater to use cases where distinguishing negatives is crucial. AdaBoost, with its top-tier specificity, is particularly suitable for conservative applications that aim to minimize false alarms, though its weak recall limits broader clinical utility.

ROC curve of top performing models

The ROC (Receiver Operating Characteristic) curve (Fig. 5) illustrates each model’s ability to distinguish between positive and negative cases by plotting sensitivity against 1-specificity. Among the top models, Random Forest showed the highest sensitivity (0.95) but low specificity, leading to more false positives. XGBoost offered a better balance with strong sensitivity (0.78) and high specificity (0.88), placing it closer to the ideal point on the ROC curve. SVM and ANN performed moderately. Overall, Random Forest and XGBoost outperformed the others, with XGBoost providing a more clinically practical balance between accurate detection and fewer false alarms.

Fig. 5
figure 5

ROC curves of top performing ML models.

While Random Forest delivers the most comprehensive performance, its low specificity may require mitigation strategies in clinical settings. SVM and LR remain dependable choices, especially when interpretability or sensitivity is prioritized. KNN, XGBoost, and AdaBoost serve niche roles based on the desired trade-offs between false positives and negatives. Ultimately, the optimal model depends on the clinical or operational context, and the inclusion of 95% confidence intervals provides a crucial statistical basis for informed decision-making.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *