The accuracy of the models used to predict the viscosity of pure and mixed imidazolium-based ILs is thoroughly evaluated using statistical errors and graphical plots. The results are presented in the following sub-sections.
Statistical error evaluation
The accuracy of the models was evaluated by calculating the difference between the predicted viscosity (ηpred) and the experimental viscosity (ηexp). To do so, five different measures, including average percent relative error (APRE), average absolute percent relative error (AAPRE), standard deviation (SD), coefficient of determination (R2), and root mean square error (RMSE), were used. The mathematical representations of these statistical indicators are presented below:
$$APRE = \frac{1}{n}\mathop \sum \limits_{i = 1}^{n} \frac{{\left( {\eta_{i,exp} – \eta_{i,pred} } \right)}}{{\left( {\eta_{i,exp} } \right)}} \times 100$$
(5)
$$AAPRE = \frac{1}{n}\mathop \sum \limits_{i = 1}^{n} \left| {\frac{{\left( {\eta_{i,exp} – \eta_{i,pred} } \right)}}{{\left( {\eta_{i,exp} } \right)}}} \right| \times 100$$
(6)
$$RMSE = \sqrt {\frac{{\mathop \sum \nolimits_{i = 1}^{n} \left( {\eta_{i,exp} – \eta_{i,pred} } \right)^{2} }}{n}}$$
(7)
$$SD = \sqrt {\frac{{\mathop \sum \nolimits_{i = 1}^{n} \frac{{\left( {\eta_{i,exp} – \eta_{i,pred} } \right)}}{{\eta_{i,exp} }}^{2} }}{n – 1}}$$
(8)
$${\varvec{R}}^{2} = 1 – \frac{{\mathop \sum \nolimits_{{{\varvec{i}} = 1}}^{{\varvec{n}}} \left( {{\varvec{\eta}}_{{{\varvec{i}},{\varvec{exp}}}} – {\varvec{\eta}}_{{{\varvec{i}},{\varvec{pred}}}} } \right)^{2} }}{{\mathop \sum \nolimits_{{{\varvec{i}} = 1}}^{{\varvec{n}}} \left( {{\varvec{\eta}}_{{{\varvec{i}},{\varvec{exp}}}} – \overline{\user2{\eta }}_{{{\varvec{i}},{\varvec{exp}}}} } \right)^{2} }}$$
(9)
Table 3 shows the statistical errors calculated for the RF, CatBoost, KNN, and LightGBM models implemented to predict the viscosity of pure ILs. The studied dataset for pure viscosity prediction consists of a training set with 3961 data points and a testing set with 991 data points. Table 4 shows the statistical errors calculated for the RF, CatBoost, GPR, and LightGBM models used to predict the viscosity of IL mixtures. The studied dataset for mixture viscosity prediction consists of a training set and a testing set with 1181 and 296 data points, respectively. Performance comparison of the pure IL viscosity prediction models, as shown in Table 3, reveals that the RF model achieved the highest overall accuracy with the lowest AAPRE and RMSE values, along with the highest R2. This indicates its strong predictive capability and excellent generalization performance. The CatBoost model performs competitively, especially during training, but shows slightly lower accuracy than RF on the test data, which may be due to the complexity associated with handling categorical variables in new samples. The LightGBM model shows the weakest performance among the models, probably due to its sensitivity to specific data features, resulting in higher errors and reduced prediction accuracy. The KNN model shows similar error statistics to the RF model because both models can capture nonlinear relationships without assuming a specific parametric form. The KNN model predicts outcomes based on the local similarity of samples and effectively models complex patterns, especially in low-dimensional spaces. However, the KNN model is more sensitive to noise and neighbor selection, and unlike the RF model, it does not aggregate information across multiple models. The ensemble of decision trees in the RF model helps reduce variance and improve robustness against noise and irrelevant features. This ensemble approach enables the RF model to better generalize and handle complex, high-dimensional data, and explains why it consistently outperforms the KNN model despite similar error metrics. The performance comparison of the models for predicting the viscosity of mixtures, as shown in Table 4, reveals that the CatBoost model achieves the best overall accuracy, with the lowest AAPRE, RMSE, and SD values, and the highest R2, indicating excellent predictive performance and generalization. The GPR model also performs well, with consistently low error metrics and high R2 values across all datasets, reflecting its strength in modeling smooth and complex nonlinear relationships. In contrast, the RF and LightGBM models exhibit higher error values and lower R2 scores than other ML models, likely due to their sensitivity to feature interactions and their limitations in capturing the complexities of mixture viscosity behavior. In contrast, the traditional models ePC-FVT-MB and ePC-FVT-IB exhibited significantly higher errors and lower R2 values, highlighting their limited flexibility and accuracy compared to data-driven ML models. Overall, ML approaches demonstrated superior performance due to their adaptability and ability to capture intricate relationships in diverse mixture compositions. Table 5 summarizes the best-performing models for predicting the viscosity of pure and mixed imidazolium-based ionic liquids across training, testing, and total datasets. The RF model showed the best performance for pure ILs, while the CatBoost model achieved the highest accuracy for mixed ILs, with the lowest AAPRE and RMSE and highest R2 in all dataset splits. This highlights the strength of RF in handling pure IL systems and the superior ability of CatBoost in capturing complex patterns in IL mixtures.
Graphical error analysis
Graphical error analysis is used as an effective technique to evaluate the performance of the models. This method is especially important for comparing the accuracy of different models. In this study, several analyses were performed to demonstrate the efficiency of the developed model. The display of different graphs, such as cross graphs, error distribution, grouped errors, and cumulative frequency graphs, was used to show the performance of the created models.
The cross-plot shows the predicted values (Pred) versus the experimental values (Exp). The straight line at an angle of 45° in this plot represents the ideal model. The higher the concentration of points around this line, the more accurate the model prediction is. Figure 2 shows the cross-plots for the different models. In Fig. 2a, the pure viscosity values predicted by the RF models lie well around the line Y = X for the training and test sets. Only a limited number of viscosity data points do not lie around the straight 45° line. Also, in Fig. 2b, for the mixture viscosity, the values predicted by the CatBoost model lie well on the line with a unit slope. In Fig. 2, the dense concentration of points around the 45° line indicates the highest accuracy for the RF and CatBoost models in predicting viscosity, thus confirming the findings presented in Tables 3 and 4. Furthermore, the high scattering of the points for the ePC-FVT-MB and ePC-FVT-IB models in Fig. 2b indicates that these two models have higher errors and lower accuracy for predicting mixture viscosity than the ML models.
Cross-plot of the models used for viscosity prediction of (a) Pure ILs, (b) IL mixtures, and c) ePC-FVT models for IL mixtures.
The error distribution curve is a statistical method used to display the error values of each model against the experimental data points. In this graph, the higher the concentration of points around the Y = 0 line, the lower the error and the higher the accuracy of the model. In this graph, the y-axis represents the relative error between the predicted and experimental data, while the x-axis represents the experimental data. According to Fig. 3a, the relative error range for all ML models is between -3 and 3, which indicates the high accuracy of all models in predicting the viscosity of pure ILs. However, among these models, the RF model has a lower relative error range than other models, and the density of points on the Y = 0 line is higher, indicating the higher accuracy of this model compared to the size of the models for predicting the viscosity of pure ILs. Also, the relative error values in terms of experimental values of mixture viscosity in Fig. 3b show that the ePC-FVT-MB and ePC-FVT-IB models have higher relative errors than ML learning models. However, among the ML models, the CatBoost model has a lower relative error range than other models, which indicates the high accuracy of this model for predicting mixture viscosity.


Error distribution plots for the proposed models in predicting the viscosity of: (a) Pure ILs, and (b) IL mixtures.
Figure 4 shows the cumulative frequency versus relative error for all models. This method charts the cumulative frequency of data points versus the relative error (at a specified threshold) to assess the number of data points the models can predict with accuracy and reliability. A higher cumulative frequency indicates a larger portion of the data set with estimation errors of a given value or less, indicating a greater ability of the model to produce reliable results. As Fig. 4a depicts, the LightGBM model has the worst performance among all models. However, the RF model outperforms other models, with more than 90% of its predictions having an error of less than 0.1. A glance at Fig. 4b reveals that the CatBoost model has a lower error than other models, predicting more than 90% of the data with a relative error of less than 0.06.

Cumulative frequency plots for the models used in this study to predict the viscosity of (a) Pure ILs, and (b) IL mixtures.
Figure 5 compares the relative errors of the studied models, calculated as the ratio of the difference between the experimental and predicted viscosity to the experimental viscosity. According to Fig. 5a, the LightGBM model exhibits the widest error range, while the KNN, CatBoost, and RF models demonstrate similar performance. Among them, the RF model shows the narrowest error distribution, ranging from -1.21 to 0.55. Figure 5b indicates that the CatBoost model has the smallest relative error range, from -0.48 to 0.22, whereas the ePC-FVT-MB model displays the widest error range and the weakest performance among all models.

Error distribution in viscosity prediction across various ML models for (a) Pure ILs, and (b) IL mixtures.
Group error plots are an effective way to evaluate the performance of the model over different ranges of input parameters. Figure 6 shows the absolute relative error over different ranges for the input parameters, as well as the absolute relative error for pure ILs and IL mixtures. Figure 6a shows the absolute relative error plot for the input parameters temperature (T), critical temperature (Tc), critical pressure (Pc), and boiling point (Tb), which are used to predict the viscosity of pure ILs. These plots provide insight into how the accuracy of the model changes with different physical and thermodynamic properties. Figure 6a shows that the RF and KNN models show similar performance, but the RF model generally achieves lower error and higher accuracy. In contrast, the LightGBM model shows the highest error and lowest accuracy among the ML models. Figure 6b shows the absolute relative error for the pure ILs used in this work. The RF model has a lower absolute relative error for most ILs than the other models. Figure 6c shows the absolute relative error graph in different ranges of input parameters: temperature (T), critical temperature (Tc), critical pressure (Pc), and acentric factor (ω) used to predict the viscosity of the IL mixtures. As shown in Fig. 6c, the CatBoost model has a lower absolute relative error than other models in all ranges of input parameters, which indicates a lower error and higher accuracy of this model than other models. Figure 6d demonstrates that the CatBoost, GPR, and RF models exhibit the lowest errors and highest accuracy among all models, while the ePC-FVT-IB and ePC-FVT-MB models show relatively higher errors and lower accuracy. Although the ePC-FVT-IB model performs better than the ePC-FVT-MB model in terms of error, the ML models, particularly CatBoost, GPR, and RF, demonstrate significantly superior accuracy. Among all models studied, the CatBoost model achieved the highest overall accuracy, making it the most reliable for predicting the viscosity of IL mixtures.


Group error diagrams of proposed models for pure ILs (a and b) and IL mixtures (c and d).
Figures 7a and b provide the relative error (%) of the RF and Cat-Boost models that proved to be the best models for predicting the viscosity of pure ILs and IL mixtures studied in this work. According to this Box and Whisker plot, the high accuracy and reliability of these models for predicting the viscosity of imidazolium-based ILs can be inferred.

Box and whisker plot showing relative error (%) for (a) pure ILs, and (b) IL mixtures.
The SHAP (SHapley Additive exPlanations) algorithm is employed in this study to assess the significance of input features in the RF and CatBoost models. Figure 8 presents summary plots that depict the influence of each parameter along with the corresponding SHAP values for each feature. The y-axis lists the input variables along with their significance, while the x-axis displays SHAP values, representing their influence on predictions. The color scale, transitioning from blue to red, indicates how variations in input variable magnitudes affect the model’s output. The SHAP (Shapley Additive Explanations) summary plot in Fig. 8a illustrates the influence of each input feature on the RF model’s viscosity predictions for pure ILs, while Fig. 8b depicts the effect of each input feature on the CatBoost model’s viscosity predictions for IL mixtures. The summary plot identifies the most influential factors in the viscosity prediction, with temperature (T) having the greatest impact, as evidenced by the wide SHAP values. The model behavior generally aligns with physical intuition, enhancing its reliability and interpretability. This analysis shows that the RF and CatBoost models effectively capture the key thermodynamic and molecular properties that affect the viscosities of pure and mixed ILs, respectively, and increase confidence in their predictive performance.

Distribution of SHAP value of each sample for viscosity prediction of a) pure ILs, and b) IL mixtures.
Model trend analysis
Figure 9a shows that the viscosity of [C8mim][PF6] at a constant pressure of 101 kPa decreases with increasing temperature, and the RF model was able to predict this relationship well, so that the value predicted by the RF model was well matched with the experimental data. Figure 9b depicts the viscosity of [C8mim][PF6] at constant temperature and varying pressures. According to this Figure, the value of viscosity increases with increasing pressure, and the RF model could predict this relationship well. However, [C8mim][PF6] viscosity also decreased with increasing pressure, demonstrating the low accuracy of the model. Figure 9c illustrates the viscosity of a mixture of 20% [C2mim][OAc] and 80% [C2mim][C2SO4] at constant pressure versus temperature. As observed, the viscosity of the mixture decreases with increasing temperature, and the data predicted by the CatBoost model are in very good agreement with the experimental data, indicating the high accuracy of the model in the entire temperature range. Figure 9d shows the viscosity of the mixture [C4mim][Tf2N] + [C4mim][SCN] at constant pressure in terms of the molar fraction of [C4mim][Tf2N] and at two constant temperatures of 338.15 K and 358.15 K. This Figure shows that with increasing [C4mim][Tf2N] concentration, its viscosity value decreases, and the CatBoost model was able to predict this relationship well. This analysis is consistent with physical intuition, because viscosity generally decreases with temperature and increases with increasing pressure. Conclusively, both RF and CatBoost models could accurately predict the viscosity of pure IL and IL mixtures.

(a) Effect of temperature change on [C8mim][PF6] viscosity at 101.325 kPa. (b) Effect of pressure change on the viscosity of [C8mim][PF6] at 353.15 K. (c) Effect of temperature change on the viscosity of a mixture of 20% [C2mim][OAc] and 80% [C2mim][C2SO4]. (d) Effect of changing the molar fraction of [C4mim][Tf2N] on the viscosity of a mixture of [C4mim][TF2N] and [C4mim][SCN] at 358.15 K and 338.15 K.
Applicability domain and outlier identification in RF and CatBoost models
Following statistical and graphical assessments, which confirmed the RF model’s effectiveness for viscosity prediction of pure ILs and the CatBoost model’s accuracy for viscosity prediction of IL mixtures, an outlier detection method was applied to identify anomalous data that might negatively impact model predictions. In this approach, data points that deviate considerably from the majority are identified as outliers. The Leverage technique is based on three key components: Standardized Residuals (SR), the Leverage threshold (H*), and the Hat matrix (H). These components are defined as follows60,61,62:
$$H=X{\left({X}^{T} X\right)}^{-1}{X}^{T}$$
(10)
$${H}^{*}=\frac{3\times (Number of features+1)}{Number of data points}$$
(11)
$$SR=\frac{{e}_{j}}{\sqrt{MSE\left(1-{H}_{j}\right)}}$$
(12)
where ej represents the ordinary residual at the jth index, MSE denotes the mean squared error, and Hj corresponds to the leverage value at the jth data point. The Leverage technique utilizes the Williams plot to visually interpret results, categorizing data points into “out of leverage,” “valid data,” and “suspected data.” Each category occupies a specific region within the plot. Data points with Hat values below the H* threshold fall within the model’s applicability domain. However, points within this range, but with SR values exceeding 3 or dropping below -3, are flagged as suspected data. Consequently, “valid data” consists of points with Hat values under H* and SR values between -3 and 3, while any point surpassing H* is considered outside the model’s applicability domain63,64,65. Examination of the possible outliers in Fig. 10a shows that the RF model has 176 out of leverage and 66 suspected data points, which account for about 3.55% and 1.62% of the data points, respectively. As a result, 4,710 data points were classified as valid, making up approximately 95.11% of the total pure viscosity dataset. Similarly, Fig. 10b illustrates that the CatBoost model identified 53 out-of-leverage points and 22 suspected data points, leaving 1,402 valid data points, which represent 94.92% of the total mixed viscosity dataset. Analyzing the Williams plot confirms that the majority of data points lie within the valid range, highlighting the reliability of the dataset and the effectiveness of the data collection methods used in this study.

Williams diagram for a) pure ILs, and b) IL mixtures, revealing that RF and CatBoost models are the most accurate models for predicting the viscosity of pure ILs and IL mixtures, respectively.
