Machine learning-based academic performance prediction with explainability for enhanced decision-making in educational institutions

Machine Learning


This section outlines the methodological framework used to predict student performance. The process began with data preprocessing to prepare the dataset for analysis. Next, nine individual ML regression models were selected. Afterward, the performance of both the voting regressor and the base regressors was evaluated and compared using three different metrics on the validation dataset. Finally, the section presents an explanation and interpretation of the model’s predictions. Figure 1 illustrates the overall architecture of the proposed student performance prediction system.

Fig. 1
figure 1

Overall architecture of the proposed student performance prediction system.

Data collection and preprocessing

There are two datasets utilized in this study. The first dataset provides a comprehensive representation of students’ academic performance by including various features such as grades, Sleep Hours, and extracurricular activities. It comprises 10,000 rows and six features14, including ‘Hours Studied’, ‘Previous Scores’, ‘Extracurricular Activities’, ‘Sleep Hours’, ‘Sample Question Papers Practiced’, and ‘Performance Index’—which serves as the target variable reflecting overall academic achievement. The dataset, originally obtained from Kaggle, has been validated to contain no missing or duplicate records, ensuring its suitability for analysis. These features are a mix of numerical and categorical variables. ‘Hours Studied’, ‘Previous Scores’, ‘Sleep Hours’, and ‘Sample Question Papers Practiced’ are continuous numerical variables, while ‘Extracurricular Activities’ is a binary categorical feature (Yes/No). As shown in Table 2, each feature is described in detail to highlight its relevance. Preprocessing is a critical step for preparing the dataset for ML models. First, categorical features such as extracurricular activities are encoded using a label encoder to convert them into numerical values. Subsequently, numerical features are normalized using a standard scaler to normalize the data distribution and enhance the models’ convergence during training. Table 3 illustrates some examples of the first dataset. Table 4 highlights the statistical analysis of the first dataset.

The second dataset is that student performance factors thoroughly examine the numerous variables that influence student exam performance15. It encompasses data regarding attendance, study practices, parental involvement, and other factors contributing to academic success. This dataset has 6,607 records and 20 features as shown in Table 5. There are seven numeric columns, including exam score, Attendance, and hours studied. There are 13 categorical columns, including gender, parental Involvement, access to resources, extracurricular activities, and teacher quality. The teacher quality, parental education level, and distance from home columns of the dataset contain missing values. The most frequently occurring value will be used to replace the missing values in each categorical column and with the mean value in numerical features. We identified 101 as an outlier value in the exam score column; therefore, we have replaced it with 100. All the categorical features are converted into numerical values. Table 6 illustrates some examples of the second dataset.

Table 2 Description of the features of the first dataset.
Table 3 Examples of the first dataset.
Table 4 The statistical analysis of the first dataset.
Table 5 Description of various features of the second dataset.
Table 6 Examples of the second dataset.

Machine learning algorithms

To predict the student academic performance, we implemented nine ML algorithms, including Linear Regression, K-Nearest Neighbors (KNN), XGBoost, CatBoost, AdaBoost, Random Forest, Support Vector Regression (SVR), Ridge Regression, and Bagging Regressor. These models were selected based on their effectiveness in regression tasks and their ability to capture diverse patterns in educational data. Below, we detail each algorithm’s role in predicting student outcomes, along with their mathematical formulations.

  1. 1)

    Multi-Linear Regression (MLR).

MLR models the relationship between a dependent variable (final grade) and multiple independent features (e.g., attendance, past grades, extracurricular activities) using a linear approach16. The model minimizes the residual sum of squares (RSS) to optimize coefficients. Mathematical formulation is:

$$\:\widehat{y}={\beta\:}_{0}+{\beta\:}_{1}{x}_{1}+{\beta\:}_{2}{x}_{2}+\dots\:+{\beta\:}_{n}{x}_{n}+\epsilon$$

(1)

.

where \(\:\widehat{y\:}\)is the predicted grade, \(\:{\beta\:}_{0}\)​ is the interception, \(\:{\beta\:}_{i}\)​ are coefficients, ​\(\:{x}_{i}\) are input features, and \(\epsilon\) is the error term. In our study, MLR identifies the influence of each feature (e.g., attendance, grades, extracurricular participation) on the final grade. This algorithm serves as a baseline for comparing other regression methods.

  1. 2)

    K-Nearest Neighbors Regression (KNN).

KNN regression is an instance-based algorithm that predicts the output for a given input by averaging the target values of the K-nearest data points in the feature space17,18. The following is determined using a distance metric, such as Euclidean distance. In our method, KNN predicts the performance of a student from the average of the grade of KK, using the Euclidean distance for proximity, the most similar students in the feature space are from the average of the same students. KNN can formally represent:

$$\:\widehat{y}\left(x\right)=\frac{1}{K}\sum\:_{i=1}^{K}{y}_{i}$$

(2)

.

where \(\:{y}_{i}\)​ are the target values of the K nearest neighbors. KNN effectively captures localized trends, making it useful for identifying peer-influenced performance patterns.

  1. 3)

    XGBoost Regression.

XGBoost, or Extreme Gradient Boosting, is an advanced implementation of promoting shields that involves regularization to prevent overfitting and increase computational efficiency19,20. It constructs decision trees iteratively, optimizing an objective function at each step. XGBoost is applied in our model to exploit its ability to capture complex, non-linear relationships among features, thereby improving prediction accuracy for diverse student data. Mathematically it can represent as:

$$\:\mathcal{L}\left(\theta\:\right)=\sum\:_{i=1}^{n}l\left({y}_{i},{\widehat{y}}_{i}\right)+\sum\:_{k=1}^{K}{\Omega\:}\left({f}_{k}\right)$$

(3)

.

where l is the loss function, Ω penalizes model complexity, and ​\(\:{f}_{k}\) represents trees. The grid search used to optimize the hyperparameters, and the optimal values, which are n_estimators = 100, max_depth = 3, learning_rate = 0.1, min_child_weight = 1, gamma = 0, subsample = 0.8.

  1. 4)

    CatBoost Regression.

CatBoost, short for “Categorical Boosting,” is a gradient boosting algorithm tailored for categorical data21. CatBoost handles categorical features (e.g., course type, extracurriculars) natively, reducing preprocessing needs and overfitting. In our case, CatBoost enhances predictions by leveraging its capability to process categorical attributes, such as extracurricular activities, directly within the model. It can be represented mathematically as:

$$\:{\widehat{y}}_{i}=\sum\:_{j=1}^{M}\eta\:.{h}_{j}\left({x}_{i}\right)$$

(4)

.

where \(\:{h}_{j\:}\)are the decision trees, and η is the learning rate.

  1. 5)

    AdaBoost Regression.

AdaBoost, or Adaptive Boosting, is an ensemble method that consecutively combines uncertain learners, such as decision trees, to enhance overall prediction accuracy22. Each subsequent learner focuses on correcting errors made by previous learners. For our model, AdaBoost captures intricate patterns in the data and provides an ensemble-based perspective on student performance trends. It can be represented as:

$$\:\widehat{y}\left(x\right)=\sum\:_{t=1}^{T}{\alpha\:}_{t}{h}_{t}\left(x\right)$$

(5)

.

where \(\:{\alpha\:}_{t}\) are weights, and \(\:{h}_{t}\) are weak learners.

  1. 6)

    Random Forest Regressor.

An ensemble of decision trees that reduces variance by averaging predictions from multiple trees23. Random Forest mitigates overfitting and handles high-dimensional educational data well. It represents as:

$$\:\widehat{y}\left(x\right)=\frac{1}{B}\sum\:_{b=1}^{B}{T}_{b}\left(x\right)$$

(6)

.

where B is the number of trees, and Tb(x) are individual tree predictions.

  1. 7)

    SVR.

SVR maps feature high-dimensional space to fit a hyperplane with maximum margin tolerance24. SVR is effective for small datasets with non-linear relationships. The mathematical formulation of it is as follows:

$$\:min\frac{1}{2}{\parallel w\parallel}^{2}+C\sum\:_{i=1}^{n}\left({\xi\:}_{i}+{\xi\:}_{i}^{*}\right)\:\text{s}\text{u}\text{b}\text{j}\text{e}\text{c}\text{t}\:\text{t}\text{o}\:\mid{y}_{i}-\langle w,{x}_{i}\rangle-b\mid\le\:\epsilon\:+{\xi\:}_{i}$$

(7)

.

  1. 8)

    Ridge Regression.

A regularized linear regression variant that prevents overfitting via L2 penalty25. Ridge improves stability in multicollinear features (e.g., correlated course grades). Mathematically represented as:

$$\:\widehat{y}=\text{arg}\underset{w}{\text{min}}{\parallel y-Xw\parallel}^{2}+\alpha\:{\parallel w\parallel}^{2}$$

(8)

.

Ridge improves stability in multicollinear features (e.g., correlated course grades).

  1. 9)

    Bagging Regressor.

Aggregates predictions from multiple base models (e.g., decision trees) trained on random subsets of data26. Bagging reduces variance and enhances generalization. It represented mathematically as:

$$\:\widehat{y}\left(x\right)=\frac{1}{M}\sum\:_{m=1}^{M}{F}_{m}\left(x\right)$$

(9)

.

  1. 10)

    Ensemble Voting Regressor (Proposed Model).

The proposed Ensemble Voting Regressor (VR) was designed to enhance prediction accuracy by strategically combining the strengths of multiple machine learning models while mitigating their individual limitations. Prior to model training, feature selection was conducted to eliminate redundant or non-predictive variables, ensuring computational efficiency without compromising performance. This process relied on correlation analysis, as visualized in Fig. 2, which depicts Pearson correlation coefficients between input features and the target variable (Performance Index) of the first dataset. The analysis revealed that Previous Scores exhibited an exceptionally strong positive correlation (r = 0.92), establishing it as the most influential predictor of academic performance. Hours Studied showed a moderate correlation (r = 0.37), while other features—such as Extracurricular Activities (r = 0.02), Sleep Hours (r = 0.05), and Sample Question Papers Practiced (r = 0.04)—demonstrated negligible relationships with the target. Consequently, only Previous Scores and Hours Studied were retained for the feature-selected subset, streamlining the model by removing noise from weakly correlated variables. Figure 3 depicts Pearson correlation coefficients between input features and the target variable (Exam Score) of the second dataset. The analysis revealed that Attendence exhibited an exceptionally strong positive correlation (r = 0.58), establishing it as the most influential predictor of exam score. Hours_Studied showed a moderate correlation (r = 0.45).

Fig. 2
figure 2

Pearson correlation heatmap showing the relationship between input features and the target (Performance Index). Features with low correlation values (e.g., Extracurricular Activities) were excluded during feature selection.

Fig. 3
figure 3

Pearson correlation heatmap showing the relationship between input features and the target (Exam Score).

The VR model operates by aggregating predictions from the top five standalone algorithms, weighted according to their individual performance. For the feature-selected scenario, the ensemble incorporated Linear Regression, Ridge Regression, CatBoost, XGBoost, and Random Forest, with weights assigned as1,7 to prioritize linear models, which outperformed others in preliminary tests. In the full-feature scenario, SVR replaced Random Forest due to its better handling of high-dimensional data, though the weighting scheme remained consistent. The weights were optimized via grid search to minimize the root mean squared error (RMSE), ensuring the ensemble’s predictions balanced bias and variance effectively. Mathematically, the VR’s prediction \(\:\widehat{y}VR\)​ is derived as a weighted sum of base model outputs:

$$\:\widehat{y}VR=\:\sum\:_{i=1}^{5}{w}_{i}.{\widehat{y}}_{i}\:\:\text{w}\text{h}\text{e}\text{r}\text{e}\:\sum\:{w}_{i}=1$$

(10)

.

where \(\:{w}_{i}\)​ represents the normalized weight of the i-th model. Figure 4 illustrates the three-stage architecture of our proposed Ensemble Voting Regressor, which strategically combines predictions from five carefully selected base models to achieve optimal performance in student academic prediction. In our implementation:

  1. 1)

    Transition Set (Base Models):

We employed five distinct regression models that demonstrated complementary strengths in preliminary testing:

  • Linear Regression and Ridge Regression (prioritized for their interpretability and stable performance on linear relationships).

  • CatBoost and XGBoost (selected for their ability to capture complex non-linear patterns).

  • Random Forest (including for its robustness against overfitting).

Each model processes the input features (notably Previous Scores and Hours Studied after feature selection) to generate independent predictions (P₁ to P₅). The diversity of these models ensures that different aspects of student performance patterns are captured, from simple linear correlations to intricate non-linear relationships.

  1. 2)

    Voting Mechanism:

Our weighted voting scheme assigns importance values of1,7 to the respective models, reflecting their relative predictive performance. This weighting:

  • Strongly favors the linear models (7× weight each) due to their superior baseline performance.

  • Moderately incorporates the tree-based models (1× weight each) to add nuanced pattern recognition.

  • It was optimized through grid search to balance the strengths of each approac.

  1. 3)

    Final Prediction:

The ensemble’s output combines these weighted predictions into a single, more accurate estimate of student performance. The architecture’s effectiveness is particularly evident in how it handles edge cases – where one model might struggle, others compensate. For instance, when predicting performance for students with unusual study patterns (very high Hours Studied but average Previous Scores), the linear models provide stability while the tree-based models adjust for non-linear effects.

To address the “black-box” nature of ensemble methods, LIME and SHAP were employed27. LIME provided localized insights by approximating the VR model’s behavior for individual predictions (e.g., identifying that a student’s low Hours Studied value disproportionately reduced their predicted grade). SHAP offered a global perspective, quantifying the average contribution of each feature across the dataset. For instance, SHAP values confirmed that Previous Scores accounted for 62% of the VR model’s output, aligning with the correlation heatmap. These interpretability tools not only validated the model’s logic but also made its predictions actionable for educators—highlighting which factors (e.g., prior academic performance) were most critical for intervention.

Fig. 4
figure 4

Architecture of the proposed Ensemble Voting Regressor.

The VR model achieved superior performance compared to all standalone algorithms, with an RMSE of 0.1050, MAE of 0.0837, and R² of 0.9890. Figure 5 further demonstrates the ensemble’s stability, showing how the voting mechanism reduces error dispersion compared to base models like SVR and Random Forest, which exhibited higher variance in predictions. This stability, coupled with the model’s transparency through LIME/SHAP, positions the VR as a reliable tool for academic advisors seeking to identify at-risk students and tailor support strategies.

Table 7 highlights the trade-offs between interpretability, data handling, and predictive power, aiding educators in selecting suitable models for academic performance analysis.

Fig. 5
figure 5

Distribution of student participation in extracurricular activities.

Table 7 Comparative summary of ML Algorithms.

Performance metrics

To evaluate the effectiveness of the proposed models, we use three standardized regression evaluation metrics28,29:

Calculates the average magnitude of prediction errors, expressed as:

$$MAE =\:\:\:\frac{1}{n}\:\sum\:_{i=1}^{n}\left|{x}_{i}-\:\left.{y}_{i}\right|\right.$$

(11)

MAE is particularly useful for interpreting the average prediction error in the same units as the dependent variable.

Quantifies the standard deviation of prediction errors and penalizes larger errors more than MAE, calculated as:

$$RMSE =\:\sqrt{\frac{1}{n}\:\sum\:_{i=1}^{n}\left({x}_{i-\:}{y}_{i}\right)}2$$

(12)

RMSE provides a sensitive measure for identifying significant deviations in predictions.

Indicates the extent to which the independent variables account for the variation observed in the dependent variable, and is determined using the formula:

$$R^2 = 1 – \frac{{SSR}}{{SST}}$$

(13)

where SSR represents the sum of squared residuals, and SST denotes the total sum of squares. Higher R² values indicate better model fit. By integrating these metrics, we ensure a comprehensive evaluation of model performance across various dimensions, facilitating reliable predictions and actionable insights.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *