Integrating CNN and transformer architectures for superior Arabic printed and handwriting characters classification

Machine Learning


This method’s primary goal is to build Arabic-printed character classification models using an Arabic handwritten dataset. Figure 3 illustrates the workflow of the proposed hybrid feature fusion ensemble-based model combined with a transformer ViT architecture for Arabic printed character classification. The methodology encompasses several sequential steps: Initially, the input data is prepared using the Arabic-printed character dataset in conjunction with the AHCR dataset. Subsequently, the images undergo preprocessing, which includes resizing the inputs. The next step involves feature extraction, utilizing transfer learning models; specifically, pre-trained transfer learning models are employed to obtain features from the input images. These models draw upon their learned representations from extensive datasets to proficiently capture spatial and contextual information pertinent to Arabic-printed characters. Then, the top-level features of transfer learning models are extracted without classification layers and then fused to combine complementary ensemble models. The fusion process enhances the robustness and diversity of the extracted features. Then, patch embedding, where the fused features are further segmented into embedded patches to align with the input requirements of the transformer ViT model. Next, the self-attention mechanism within the ViT processes the embedded patches to capture global dependencies and relationships between patches. The output is passed through an MLP to refine the predictions and generate classification outputs. Finally, the processed features are used to classify Arabic-printed characters effectively. This hybrid approach leverages both transfer learning for robust feature extraction and the ViT model for its powerful self-attention mechanism, resulting in improved accuracy and performance for Arabic character recognition.

Fig. 3
figure 3

The proposed hybrid features a fusion ensemble-based with a transformer ViT model for Arabic printed characters classification.

Datasets description

In this study, two datasets containing printed and handwritten characters are used. Both datasets have multiple diverse ranges of character variations per letter in multiple forms, which are usable and reusable character sets and contribute to the robustness of the model.

The Arabic char 4k OCR dataset

It is an extensive compilation of images of Arabic text, both printed and handwritten, created for training, assessing, and validating optical character recognition systems31. This dataset specifically tackles the distinctive challenges posed by the Arabic script, such as its cursive writing style, the presence of diacritical marks, and intricate ligatures. It generally features scanned or digitally produced images of printed Arabic materials, showcasing a variety of fonts with different styles, sizes, and weights. Additionally, the dataset encompasses a wide range of handwritten texts gathered from numerous writers to reflect individual variations. It includes isolated representations of Arabic characters (29 letters) in their initial, medial, final, and standalone forms with 127,027 images, as well as frequently occurring Arabic ligatures like “لا” (Lam-Alif) to enhance the overall robustness of the model.

Arabic handwritten character recognition dataset (AHCR)

AHCR is a collection of handwritten Arabic characters, typically used for developing and evaluating models for Arabic script recognition32. This dataset contains a wide range of handwritten characters, often including variations in style, size, and stroke patterns, reflecting the natural diversity in handwriting across different individuals. The AHCR dataset typically includes characters from the Arabic alphabet, which consists of 28 primary letters, and is used in this work as an ablation study to test the effectiveness of the proposed model.

Datasets splitting

Table 1 illustrates the division of the Arabic OCR dataset and the AHCR dataset into three distinct subsets: training, validation, and testing. This division for the first dataset is executed according to a specified ratio of 80% for training, 10% for validation, and 10% for testing. The training set has 101,610 images used for training the model. The validation set contains 12,716 images for monitoring performance and fine-tuning. The testing set includes 12,701 images used for an unbiased final evaluation of model performance. For the second dataset, 13,261 samples for training, 1668 samples for validation, and 1662 samples for testing.

Table 1 Splitting the Arabic characters datasets (80%, 10%, and 10%) for training, validation, and testing sets.

This splitting strategy ensures a balanced distribution across all subsets, maintaining consistency in label representation while enabling robust model training, validation, and evaluation. For preprocessing, resizing the images to a fixed dimension of 224 × 224 pixels and normalizing the pixel values to a range of [0, 1].

Artificial intelligence models

This section outlines the various artificial intelligence models employed in the methodology for Arabic printed character classification, highlighting their unique architectures, roles, and contributions to the proposed system.

Convolutional neural network model

A CNN is a feed-forward neural network designed to extract hierarchical features from input data, such as images or sounds. Utilizing backpropagation for training, CNNs can learn complex nonlinear mappings and automatically detect prominent features, demonstrating robustness to variations and distortions in input data33. In this study, we developed two separate transfer learning models, which were trained individually before their features were fused, as described below:

VGG16 model

The VGG16 model, developed by Oxford University, is a pre-trained deep learning architecture often applied through transfer learning. By leveraging its training on the ImageNet dataset, this model extracts meaningful features for similar tasks34. In this study, layers below 17 were fine-tuned and set as untrainable, optimizing its feature extraction capability.

ResNet50 model

ResNet50, a deep CNN architecture introduced by Microsoft, employs residual connections to address the vanishing gradient problem in deep networks. These connections enable the network to skip layers, improving gradient flow35. For this work, layers below 123 are set as untrainable, ensuring efficient feature extraction without overfitting.

Feature extraction

Both VGG16 and ResNet50 are utilized for the extraction of significant features from the input images depicting Arabic characters. These models, which have undergone training on extensive datasets such as ImageNet, are employed for their proficiency in recognizing fundamental spatial features, including edges, textures, and elementary patterns. In the context of this task, the models are fine-tuned by freezing the lower layers while permitting only the upper layers to acquire features pertinent to the recognition of Arabic characters.

The proposed ensemble learning model

An ensemble model is a machine-learning technique that enhances prediction accuracy by combining multiple base models. This approach aims to improve the model’s overall stability and ability to generalize by leveraging the collective strengths of the integrated models, effectively reducing the risk of overfitting. Ensemble models often outperform individual models, providing more reliable and precise predictions. As a result, they are widely used in fields such as image recognition, natural language processing, and predictive analytics36. In this study, the proposed feature extraction method adopts an ensemble strategy that merges two pre-existing deep learning models while excluding their classification layers.

This study employed an ensemble learning approach by identifying an optimal combination of pre-trained deep learning models to serve as the primary framework for feature extraction. As illustrated in Fig. 4, the ensemble method proposed herein integrates multiple pre-trained deep learning models to facilitate the feature extraction process.

Fig. 4
figure 4

The proposed hybrid features a fusion ensemble-based with a transformer ViT model for Arabic printed characters classification.

Feature fusion

After extracting features using both VGG16 and ResNet50, the next step involves fusion. The outputs from the two models are concatenated, combining their extracted features into a unified representation. This ensemble approach allows the model to benefit from the strengths of both architectures, enabling it to capture a broader range of information, from simple patterns (VGG16) to more complex representations (ResNet50). The fused features are then prepared for further processing in the subsequent stages.

The proposed hybrid transformer encoder-based model

The encoder is one of the two primary components in transformer models. Encoders are built up through several layers, with each layer composed of two parallel network channels: the Multi-Headed Self-Attention and the Feedforward Neural Network with output neurons. Encoders are the models of choice when dealing with input sequences and can function independently of one another. The input sequence is embedded from a 1-dimensional input space to a conceptual D-dimensional space. Each pair of encoders uses identical illustrations37. The Multi-Headed Self-Attention uses a scaled dot-product attention weighted sum function to associate within the input vectors and create a context matrix. The context matrix is the focal point of the encoder and is controlled by a separate layer normalization, then it is forwarded to the next part. Positional encoding can be applied pre- or post-the input sequence embedding unit. This is a vital part of the transformer model that determines the natural order of the input sequence and helps the model to understand context38.

There are two important roles played by the encoder. The first is that it is used in the training loop to minimize the discrepancy between generated probability distributions and oracle probability distributions. Second, it is used during the testing loop to decode given data. Each of the inputs is processed through four layers, from top to bottom, and then pooled to develop a joint tensor by transforming the learned representations. Since the transformer model uses self-attention, it allows the processing input vectors in parallel, thereby enabling it to compute the representations of any input sequence containing vectors in constant time.

In this paper, we propose a transformer encoder model for the task of Arabic printed characters classification. Our proposed model has a transformer encoder-based architecture that has unique features and modifications. The encoded architecture consists of a self-attention mechanism complemented by a multilayer perceptron (MLP) component. Within each block, there is an integration of a normalization layer in conjunction with residual connections. The multilayer perceptron represents a distinct variant of feed-forward neural networks characterized by the incorporation of dense layers and dropout layers, as detailed in the corresponding equations:

$$Attention\left(Q,K,V\right)=Softmax \left(\frac{Q{k}^{t}}{\sqrt{{d}_{k}}} \right)v,$$

(1)

Where Q denotes the query vector, V is the value vector with its respective dimensions, and K indicates the key vector. The product exhibits a variance that has a mean of zero. Furthermore, the product is normalized by dividing it by the standard deviation. The SoftMax function then transforms this scaled dot-product into an attention score.

The multi-head attention linearly extends the queries, keys, and values h times using a variety of learned linear projections, and can be calculated by Eqs. (2, 3).

$$MultiHead \left(Q, K, V\right)=Concat\left({head}_{1}, \dots {head}_{h}\right){W}^{o}$$

(2)

$${head}_{i}=Attention \left({QW}_{i}^{Q}, {KW}_{i}^{K}, {VW}_{i}^{V}\right),$$

(3)

the projections are parameter matrices \({\text{W}}_{i}^{Q}\in {\text{R}}^{{d}_{model} x {d}_{k}}, {\text{W}}_{i}^{K}\in {\text{R}}^{{d}_{model} x {d}_{k}}\), \({\text{W}}_{i}^{V}\in {\text{R}}^{{d}_{model} x {d}_{v}}\) and \({W}^{o}\in {\text{R}}^{{hd}_{v} x {d}_{model}}\). Conversely, the MLP component comprised a non-linear layer utilizing the Gaussian Error Linear Unit (GELU) activation function, incorporating 1024 neurons and batch normalization, with a dropout rate of 50% applied uniformly across all dropout layers.

This mechanism is fundamental to the transformer model, enabling it to provide parallel attention for analyzing the complete content of the input image. Due to the multi-head attention feature, the model can simultaneously engage with input from multiple representation subspaces located in different areas.

Transformer ViT integration

Once the features are fused, they are segmented into embedded patches to align with the input requirements of the ViT. The transformer model processes these patches using its self-attention mechanism to capture both local and global dependencies across the entire image. This hybrid approach ensures that the model can learn both fine-grained details (from CNN features) and contextual information (from ViT’s global relationships), resulting in superior performance for Arabic character classification.

By combining the strengths of transfer learning models and transformer architectures, this methodology achieves improved accuracy and robust feature extraction for Arabic printed character recognition.

Final classification

The transformer encoder processes the embedded patches, and the final classification is carried out through an MLP, which refines the feature representations and generates the classification output.

Evaluation

To evaluate the performance of deep learning models, various metrics are calculated using the outcomes of predictions: True Positives (TP), False Positives (FP), True Negatives (TN), and False Negatives (FN). The equations for the most commonly used metrics are as follows:

Accuracy:

Accuracy measures the proportion of correctly predicted samples out of the total predictions.

$$Accuracy\left( {ACC} \right)\, = \,(TP\, + \,TN) \, / \, (TP\, + \,FP\, + \,TN\, + \,FN)$$

(4)

Precision (Positive Predictive Value):

Precision indicates how many of the predicted positive instances are correct.

$$Precision\left( {PRE} \right)\, = \,TP \, / \, (TP\, + \,FP)$$

(5)

Recall (Sensitivity or True Positive Rate):

Recall measures the ability of the model to correctly identify all relevant instances.

$$Recall\left( {REC} \right)\, = \,TP \, / \, (TP\, + \,FN)$$

(6)

F1 Score:

The F1 Score is the harmonic mean of precision and recall, providing a balanced measure when precision and recall are equally important.

$$F1Score\, = \,2 \, * \, (PRE\, + \,REC) \, / \, \left( {PRE \, * \, REC} \right)$$

(7)

These metrics collectively provide a comprehensive evaluation of a model’s performance, balancing accuracy, precision, recall, and error rates to ensure a thorough analysis of its strengths and weaknesses.

The multi-class confusion metrics are utilized to obtain the values of TP, TN, FP, and FN parameters.

In addition, the area under the receiver operating characteristic (ROC) curve (AUC) serves as a metric for assessing a classifier’s ability to differentiate between distinct classes. For the multi-classification scenario, we adopt a one-class-versus-others strategy to generate the ROC curves along with their corresponding AUC values39. Subsequently, we provide the computed mean AUC values for further interpretation.

Experimental setup

In this section, we describe the experimental setup used for training the model in our study in Table 2. We focus on the configuration of key hyperparameters and the setup for the model’s training process, such as image size, batch size, learning rate, optimizer, early stopping patience, epochs, clip value, and the number of heads in multi-head attention mechanisms. Head value determines how many separate attention mechanisms the model uses to capture different relationships within the input data.

Table 2 Hyperparameter settings for the experimental setup.

For this training of the proposed AI hybrid model, it used a learning rate of 0.0001, the Adam optimizer with a clip value of 0.2, and an early stopping patience of 11. All models were trained for 25 epochs to fine-tune their hyperparameters. In the encoder section, the image size is set to 224, the patch size to 2, the input size to 20, the dropout rate for all layers to 0.01, and 8 attention heads. For handling multi-class classification tasks, the categorical cross-entropy loss function was applied.

The study utilized an MSI GS66 laptop featuring specific technical specifications, including an 11th-generation Intel Core i7 (11800H) processor, 32 GB of RAM, 2 TB of NVMe SSD storage, and an RTX 3080 graphics card equipped with 16 GB of memory. The analysis presented in this article was performed using Python 3.10 on a Windows 11 operating system, leveraging the Keras and TensorFlow backend libraries for computational tasks.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *