Sentiment analysis with echo state network and augmented water cycle algorithm

Machine Learning


Validation of the AWCA

A thorough assessment is conducted to evaluate the effectiveness of the Augmented Water Cycle Algorithm (AWCA) through a validation study using various statistical benchmark functions. These functions are commonly employed in research and concentrate on both unimodal and multimodal functions. The purpose of this validation is to assess how well different algorithms perform in solving a range of issues. Standard deviation (StD) and Average accuracy (Avg) have been employed as metrics for evaluating their effectiveness. The primary emphasis is on the impact these algorithms have on different functions within benchmark problems, where the best solution is usually located at the starting point. The stability index of the algorithms acts as an indicator of their reliability. This study employs a series of benchmark functions, F1 to F10. The findings are contextualized by comparing the results generated by the suggested algorithm and its specific metrics with those derived from four well-established algorithms known for their high reliability, including Artificial Ecosystem-based Optimization (AEO)36, Butterfly Optimization Algorithm (BOA)37, War Strategy Optimization (WSO)38, and Equilibrium Optimizer (EO)39. Table 4 shows the values of the variables used in the evaluation of each algorithm.

Table 4 The variables used in the evaluation of each algorithm.

This study thoroughly examines the suggested method against various other algorithms. Figure 2 illustrates the assessment of the AWC Algorithm in relation to different optimization techniques for enhancing the Echo State network.

Fig. 2
figure 2

The results of AWC Algorithm in comparison with other algorithms.

The results show that the AWC Algorithm has demonstrated significant effectiveness for unimodal functions (f1 to f5), which often has performed better than other algorithms regarding Average Values (AVG). This highlights its strong ability to efficiently address functions that have different global optimal points. Furthermore, for multimodal functions (f6 to f10), the AWC Algorithm performs effectively, which indicates its capability to explore and manage functions with numerous local minima successfully.

The AWC Algorithm shows superior or similar average outcomes for the majority of functions when contrasted with other algorithms. This indicates that the AWC Algorithm has successfully reached optimal or near-optimal solutions for various functions. Regarding standard deviation, while there might be some fluctuations, it demonstrates low variability or competitive in comparison with other algorithms.

The AWC Algorithm consistently produces optimal results or outcomes with minimal differences through simulations, as indicated by its low or zero standard deviation (STD) and average (AVG) values for numerous functions. The results reveal that the AWC Algorithm is an efficient optimization algorithm, which indicates a strong ability for both exploitation and exploration across numerous functions. When assessed against the other algorithms in the study, its performance, as demonstrated by the AVG and STD values, remains highly competitive.

Ablation study of the preprocessing

In this section, the purpose is to conduct an ablation study, since the effect of preprocessing is not clear. Moreover, to specify and determine the effect and strength of each preprocessing stage on the result, an ablation study has been conducted, where each preprocessing stage has been removed one each run in order to determine the efficacy of each one. The following table fully represents the results of the preprocessing stage and effect of each one (Table 5).

Table 5 Ablation study of the preprocessing.

This ablation study assesses how each preprocessing stage affects the ESN/AWCA model using Word2Vec and GloVe embeddings. The baseline, which employed all preprocessing steps, achieved the highest F1-scores of 96.37% for Word2Vec and 96.23% for (GloVe), validating that the complete process ensures optimum performance.

Among all the steps, tokenization and the removal of unwanted characters had the most substantial effect. Their absence resulted in the most substantial decline in performance, underscoring their essential role in ensuring clean and organized input. Slang correction and stopword removal produced moderate reduction in performance, indicating they are beneficial but not critical. Stemming and POS tagging showed minimal impact, reflecting their limited contribution. Generally, although tokenization and noise removal were really influential, all the stages should be conducted to gain the best results.

Simulation results

It was discussed before that diverse metrics were utilized for representing the efficiency of the recommended model compared to other strategies. Of course, these evaluation metrics represent the extent to which worked efficaciously. The findings of the performance measurement, such as precision, F1-score, accuracy, and recall, have been gained for the BiLSTM/MultiBERT6L40, ESN, BiLSTM/ELMo41, K-NN, and the suggested model ESN/AWCA on the basis of GloVe and Word2Vec. At first, the desire is to represent that if the optimization algorithm had any effect on enhancement of the findings. To this end, all the findings with the optimization algorithm and without the optimization algorithm have been illustrated below.

It has been witnessed from the findings represented in Table 6 that the optimization algorithm significantly influenced the findings. In other words, the differences in the findings of the suggested model with and without optimization algorithm have been represented. Similarly, it can be interpreted from the findings that it can generalize to other models, ensuring the influences of optimization algorithm. Therefore, this guarantees that optimization algorithm has been considered a beneficial tool for improving the results, which can be used in most of the fields. The error can be determined by subtracting 100 from the accuracy. As a result, the error amount of ESN was 13.63 while using Word2Vec, and the error amount of ESN while using GloVe was 14.86, ensuring that there is not that difference between the accuracy rate of GloVe and Word2Vec. In addition, the error amounts of ESN/AWCA were 3.63 and 3.88 while using Word2Vec and GloVe. The findings gained from the suggested model with optimization algorithm and without it have been represented using Fig. 3.

Table 6 The results of the suggested model without and with optimization algorithm while using glove and Word2Vec.
Fig. 3
figure 3

The results achieved by ESN and ESN/AWCA while using (A) Word2Vec and (B) GloVe.

The efficiency of the ESN model with and without optimization algorithm has been represented previously; however, the main purpose of the current article is to compare the efficiency of the suggested ESN/AWCA with other models and to check if this model works superior to other models or not. The ability of different neural networks, including the suggested model and the other models mentioned previously, have been compared with each other. In fact, this comparison has been conducted in order to represent the extent to which the suggested model could outperform the other models regarding their ability in categorizing the polarity of each text, whether negative or positive. It has been stated that the suggested article was carried out by the use of IMDb reviews dataset to contrast diverse models, like BiLSTM/MultiBERT6L, ESN, BiLSTM/ELMo, K-NN, and the suggested model ESN/AWCA. The findings represented that the embedding techniques GloVe and Word2Vec were really favorable that could gain superior outcomes. The efficiency of the suggested model and other models have been represented in Table 7.

Table 7 The outcomes of different models utilizing glove and Word2Vec.

On the basis of the findings accomplished and illustrated in Table 7, it can be interpreted that the performance metrics of both embedding techniques, including GloVe and Word2Vec, were nearly identical. In other words, these techniques somewhat act the same but have some little differences. However, this generalizes to other findings gained by other models as well, even those with the least proficiency and efficiency.

Initially, the main aim is to compare each model with its counterpart while using 2 diverse embedding techniques, including GloVe and Word2Vec. First, K-NN is a classical model that could not perform really well in comparison with other models. Its performance was slightly better while using GloVe with the values of 87.24%, 87.94%, 88.79%, and 87.59% for precision, recall, accuracy, and F1-score while using GloVe and could achieve the values of 87.12%, 87.65%, 87.85%, and 87.39% while using Word2Vec, respectively. It can be seen that there is not that difference between them. In the following there is another model called BiLSTM/ELMo, which is a combination of two different models. Considering this model, it could achieve the values of 91.51%, 92.31%, 91.74%, and 91.28% while using Word2Vec; however, this model accomplished slightly lower values while using GloVe. In the following, BiLSTM/MultiBERT6L was the next model that was assessed. This model could achieve better result for recall while using Word2Vec compared to GloVe with minor difference of 0.04%; however, for the rest, the model acted better while using GloVe with the minor difference of 0.34%, 0.11%, and 0.15% for precision, accuracy, and F1-score, respectively. The subsequent model is ESN without optimization algorithm that has been utilized in the present algorithm to demonstrate the efficacy and benefit of using optimization algorithm in improving the results and efficacy of a model. In this case, this network could perform better while using GloVe in all metrics compared to Word2Vec. Furthermore, the ultimate network is the one that has been suggested in the present article. In all metrics, the suggested ESN/AWCA could achieve superior results while using GloVe to Word2Vec with the minor differences of 0.14%, 0.27%, 0.11%, 0.16% for F1-score, accuracy, recall, and precision, respectively. In general, it can be seen that GloVe was somewhat better than its counterpart Word2Vec.

The LSTM model exhibited moderate performance, attaining precision, recall, accuracy, and F1-score values in the range of 89–90%, with slightly improved results when utilizing Word2Vec embeddings rather than GloVe. This suggests that LSTM is capable of effectively processing sequential data for sentiment analysis, although it did not exceed the performance of the proposed ESN/AWCA or BiLSTM/ELMo models. Conversely, the BERT model showed relatively lower outcomes in this analysis, performing similarly to the ESN without optimization and falling below both LSTM and BiLSTM-based models. This could be affected by variables such as the size of the dataset, fine-tuning techniques, or specific domain characteristics. Nevertheless, the inclusion of BERT as a baseline highlights the competitive advantage and originality of the proposed ESN/AWCA method.

Eventually, a transformer-based model has been used and optimized that has been compared with other models. Initially, it has to be mentioned that BERT/AWCA had some improvement in comparison with BERT, where BERT/AWCA could have achieved the values of 92.35, 92.62, 92.74, and 92.48 for precision, recall, accuracy, and F1-score, respectively in terms of GloVe. Moreover, that model achieved the values of 92.19, 92.31, 92.58, and 92.25 for precision, recall, accuracy, and F1-score, respectively for Word2Vec. Considering the differences in results while using the GloVe and Word2Vec, it can be observed that the model could have achieved better results in precision, recall, and F1-score while GloVe. However, the use of Word2Vec resulted in better findings in accuracy. Generally, this model had improvement while using AWCA, but it still could not reach the performance of the suggested model.

To supplement the standard evaluation metrics, ROC curves have been created to assess all models. The ESN/AWCA model gained the highest AUC of 0.973, demonstrating excellent discriminative ability. Then, BERT/AWCA and BiLSTM/ELMo achieved AUCs of 0.926 and 0.915, while the conventional ESN and BERT models had lower AUC scores of 0.865 and 0.869, respectively. These findings underscored the effectiveness of the AWCA optimizer in enhancing model robustness and accuracy across various thresholds. The results can be visualized in the following (Fig. 4).

Fig. 4
figure 4

To analyze the findings more effectively, a confusion matrix was created for the ESN/AWCA model. As illustrated in Fig. 5, the model accurately identified 11,983 out of 12,500 positive instances and 12,113 out of 12,500 negative instances. Only 517 false negatives and 387 false positives were noted. These values demonstrate the model’s high recall and precision, emphasizing its reliability in recognizing sentiment with minimal misclassification.

Fig. 5
figure 5

Generally, the best model that could achieve the highest values was, undoubtedly, ESN/AWCA. Besides, after ESN/AWCA, BERT-AWCA was the network that was the second best network among them. On the other hand, the worst network that could achieve the lowest values was ESN. These findings have been displayed in Fig. 6.

Fig. 6
figure 6

The findings gained by the recommended model and other techniques utilizing various word embedding models.

The outcomes have been presented using F1-score, precision, accuracy, and recall. The proposed sentiment analysis method surpassed the other networks regarding accuracy. This improvement was achieved by the model’s data cleaning process, which eliminated unwanted characters and corrected slang and abbreviations. Furthermore, various data preprocessing steps, such as tokenization, stemming, and part-of-speech tagging, were carried out, ultimately enhancing accuracy. In addition, during the sentiment analysis phase, these steps assisted in identifying the correct polarities of texts, resulting in increased effectiveness. However, it should be mentioned that the optimization algorithm had a key role in improving the efficacy and the results as well.

The performance of the proposed model using Word2Vec exhibited minor differences when compared to the outcomes achieved by GloVe considering the four metrics analyzed. Nevertheless, the discrepancies were not significant. The findings validated that the proposed model performed exceptionally well. The impressive results in F1-score, precision, recall, and accuracy suggested that the model’s effectiveness was improved by incorporating word embedding, comprehensive preprocessing steps, part of speech tagging, and the application of an optimization algorithm.

Validation of the results

To thoroughly verify the performance claims, a two-phase statistical analysis has been carried out. Initially, Wilcoxon signed-rank tests has been performed on the prediction results from all 25,000 test samples, comparing ESN-AWCA to other baseline models using both GloVe and Word2Vec embeddings. These non-parametric tests were chosen for their resilience to non-normal distributions and their ability to detect consistent performance differences. Subsequently, Cohen’s d effect sizes have been computed to measure the extent of improvements beyond statistical significance.

The Wilcoxon tests affirmed that the proposed ESN-AWCA significantly outperformed all baseline models. ESN-AWCA exhibited statistically greater accuracy compared to BiLSTM/ELMo (z = 5.12, p < 0.001), BiLSTM/MultiBERT6L (z = 5.87, p < 0.001), ESN (z = 6.45, p < 0.001), K-NN (z = 6.92, p < 0.001), LSTM (z = 6.15, p < 0.001), and BERT (z = 7.01, p < 0.001). Similar significant outcomes (all p < 0.001) were observed using Word2Vec embeddings, indicating consistent superiority irrespective of the embedding method. The statistical validation of the superiority of ESN-AWCA is illustrated in the table below (Table 8).

Table 8 The statistical validation of the superiority of ESN-AWCA.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *