An effectiveness of deep learning with fox optimizer-based feature selection model for securing cyberattack detection in IoT environments

Machine Learning


This article proposes the FOFSDL-SCD method. This paper analyses cybersecurity-driven techniques for improving IoT networks’ resilience and threat detection abilities utilizing advanced methods. Data pre-processing, dimensionality reduction using FOX, classification, and parameter tuning are required. Figure 2 indicates the entire workflow of the FOFSDL-SCD model.

Fig. 2
figure 2

Entire workflow of the FOFSDL-SCD method.

Pre-processing through normalization

At the primary step, the data pre-processing stage utilizes the min-max normalization method to transform the input data into a beneficial system. Min-max normalization is a data scaling method employed to convert features to a secure range, usually [0,1], conserving the relationships between original data values29. Normalization safeguards uniformity across features in cybersecurity-driven IoT networks, where network traffic and sensor data frequently differ extensively in scale. This aids ML methods in noticing anomalies more efficiently by decreasing bias toward features with greater arithmetical ranges. Applying this normalization improves the precision and consistency of intrusion detection methods in IoT systems. The mathematical formulation is given below in Eq. (1).

$$\:Y=\frac{X-{X}_{\text{m}\text{i}\text{n}}}{{X}_{\text{m}\text{a}\text{x}}-{X}_{\text{m}\text{i}\text{n}}}$$

(1)

Here, \(\:X\) signifies the original data, \(\:Y\) epitomizes the normalized data, and \(\:{X}_{\text{m}\text{a}\text{x}}\) and \(\:{X}_{\text{m}\text{i}\text{n}}\) denote the maximum and minimum values, respectively.

FOA-based feature selection method

Besides, the proposed FOFSDL-SCD designs FOA for the FS procedure to select the most significant features from the dataset. FOX is a new optimizer model stimulated by the red fox’s predatory behaviour30. The FOX is selected according to its progressive mechanisms, which deal with the restrictions of present swarm-based methods and achieve an improved balance between the exploitation and exploration stages. This mathematically verified performance highlights FOX’s adaptability, applicability, and robustness to composite optimizer tasks containing reinforcement learning, whereas practical exploitation and exploration are essential. Still, FOX intends to recognize the best solution by assessing the optimal fitness values through the searching agent’s population. The FOX works over numerous iterations with many agents, all searching for the best value of fitness (improved solution). It includes dual main steps: exploration and exploitation. During the exploration step, agents use the random walking approach to find possible solutions for how red foxes look for their victims. They utilize their capability to identify ultrasound signals to help with this hunt. During this stage, the agents approximate the distance to their target depending on the time the ultrasound signal takes to reach them. Therefore, FOX uses a different model for measurements, whereas agents jump after their approximation; they might take the prey according to the lapse time of the sound signal. Therefore, the agent’s success in taking prey is carefully related to its capability to understand the sound’s travelling time while jumping precisely.

$$\:s=BestPositio{n}_{i}\times\:{T}_{i}$$

(2)

The distance of FA from its prey (DFP) is described as demonstrated:

$$\:DF{P}_{i}=DS{F}_{i}\times\:0.5$$

(3)

After the FA fixes the distance to its target, it should instantly jump to capture it. The requisite jumping height was calculated utilizing the succeeding Eq. (4):

$$\:jum{p}_{i}=0.5\times\:9.81\times\:{t}^{2}$$

(4)

The FA travels to a novel location in the exploitation and exploration phases. The novel location is verified during the exploitation phase using the following Eqs. (5) and (6):

$$\:{X}_{i+1}=DF{P}_{i}\times\:jum{p}_{i}\times\:{c}_{1}$$

(5)

$$\:{X}_{i+1}=DF{P}_{i}\times\:jum{p}_{i}\times\:{c}_{2}$$

(6)

The \(\:{c}_{1}\) and \(\:{c}_{2}\) values are 0.180 0.820, individually estimated according to the jumping dynamics of the \(\:FA\). The jump agents must move both toward the opposite or the northeast direction. For the exploration phase, the novel location is computed utilizing the succeeding Eq. (7):

$$\:{X}_{i+1}=Best{X}_{i}\times\:rand\left(1,\:dim\right)\times\:\text{m}\text{i}\text{n}\left(tt\right)\times\:a$$

(7)

Now, \(\:tt\) refers to the average time, equivalent to the vector amount \(\:T\) divided by the problem size. The FOX uses a stationary trade-off approach, balancing the exploitation and exploration phases at each 0.5. During the exploration stage, the optimizer imitates the detection abilities of the fox’s target over random walks. This model mimics how the fox can seek a target in its surroundings, allowing the optimizer to discover possible solutions. This stage intends to converge on the best solution using the most positive regions recognized in the exploration stage.

The fitness function (FF) determines the classification precision and the chosen feature amounts. It maximizes the classification precision and lowers the set size of the designated attributes. Then, the succeeding FF is applied to evaluate individual solutions, as presented in Eq. (8).

$$\:Fitness=\alpha\:*\:ErrorRate+\left(1-\alpha\:\right)*\frac{\#SF}{\#All\_F}$$

(8)

Here, \(\:ErrorRate\) denotes the classification \(\:ErrorRate\) using the chosen features. \(\:ErrorRate\) is measured as the incorrect percentage classified to the number of classifications completed, specified as the value among (0,1). \(\:\#SF\) represents selected feature counts, and \(\:\#All\_F\) is the comprehensive quantity of features in the new dataset. \(\:\alpha\:\) is utilized to control the significance of subset length and classification quality.

Classification using the TCN model

The TCN method is deployed for the classification process. TCN addresses the task of acquiring either local or long-term dependency in the network data by employing dilated and causal convolution31. Various recurring techniques depend upon sequential processing, while the TCN utilizes causal convolution to maintain the temporal data order, guaranteeing that upcoming data is not employed to forecast preceding events. Furthermore, dilated convolution increases the receptive area without rising parameter counts, allowing the technique to acquire longer‐range dependencies more effectively. This makes TCN efficient in handling either long‐term patterns or short‐term variations. This integration of dilated and causal convolutions aids in enhancing the performance of the model. The convolutional module successfully takes the local time dependency in the input data over a convolutional operation. Conventional RNN contains superior computational efficacy and enhanced parallelization proficiencies while processing long time-series data. This model comprises dilated and causal convolutions and residual links. Dilated convolution increases the receptive area without dropping resolution; causal convolution safeguards the technique about the temporal sequence of the data, and residual connection aids in reducing the issue of gradient vanishing. Assume that the sequence of input \(\:Z=[{z}_{1},{z}_{2},\:\dots\:,{z}_{n}]\); here, \(\:{z}_{i}\) refers to the output of the self‐attention module.

$$\:{Y}_{t}={\sum\:}_{k=0}^{K-1}{W}_{k}\cdot\:{Z}_{t-k}$$

(9)

\(\:{W}_{k}\) signifies convolution kernel weight, \(\:K\) represents the size of the convolution kernel, and \(\:{Y}_{t}\) is an output of \(\:tth\) time-step. Over causal convolution, only input data of \(\:t\) is guaranteed to be employed at every time-step \(\:t\).

This method enables the technique to deal with long-time dependency by improving the parameter counts:

$$\:{Y}_{t}={\sum\:}_{k=0}^{K-1}{W}_{k}\cdot\:{Z}_{t-d\cdot\:k}$$

(10)

\(\:d\) refers to the dilation coefficient regulating the convolution kernel’s hole size. This model employs residual links to mitigate the issues of gradient explosion and disappearance. This method permits the technique to bypass definite layers, assisting the data flow and gradient.

$$\:{Y}_{t}=F\left({Z}_{t}\right)+{Z}_{t}$$

(11)

\(\:{Z}_{t}\) refers to the input, and \(\:F\left({Z}_{t}\right)\) indicates the output after dilated and causal convolution.

The convolutional time-series module is formed by many dilated and causal convolutional layers with every residual connection. This loaded framework acquires layer-wise dependencies of diverse time scales, thus enhancing the modelling proficiency of longer time-series data. Letting an output of the \(\:lth\) layer is \(\:{Y}^{l}\). The mathematical model is:

$$\:{Y}^{l}=ReLU\left(LN\left({F}^{l}\left({Z}^{l}\right)\right)\right)+{Z}^{l}$$

(12)

Where \(\:LN\) represents layer normalization, \(\:ReLU\) refers to the activation function, \(\:{Z}^{l}\) signifies the input of this layer, and \(\:{F}^{l}\) indicates the convolution operation of the \(\:lth\) layer. Within the previous layer of convolutional modelling, a fully connected (FC) layer can be employed to modify an output of the convolution process into a last value.

$$\:\widehat{Y}=OW+b$$

(13)

Now, \(\:b\) indicates a biased term, and \(\:W\) represents a weight matrix. The FC layer could map an output of a higher-dimensional convolution to the needed dimension for attaining the outcome. Over these operations, this approach can effectively take a long time and local dependencies in input data. Concurrently, the generalizability of the training method is further enhanced by over-optimization approaches like early stopping, dropout, and batch normalization. The output is integrated to present an effective and comprehensive representation. Figure 3 illustrates the framework of the TCN method.

Fig. 3
figure 3

Architecture of the TCN method.

DBO-based hyperparameter selection approach

Finally, the DBO-based hyperparameter selection method is implemented to improve the classification outcomes of the TCN model. The elementary DBO model is naturally stimulated by the foraging, dancing, stealing, breeding, and rolling behaviours of DBs32. Based on these behaviours, four population-updated tactics are designed.

Rollerball dung beetles

Naturally, DBs utilize solar navigation to manage a straight route while rolling their dung balls. Equation (14) is applied to change the location of the rolling DB:

$$\:\begin{array}{c}{x}_{i}(t+1)={x}_{i}\left(t\right)+\alpha\:\times\:k\times\:{x}_{i}(t-1)+b\times\:\varDelta\:x\\\:\varDelta\:x=\left|{x}_{i}\left(t\right)-{X}^{\omega\:}\right|\end{array}$$

(14)

Whereas \(\:t\) characterizes the present iteration counts, and \(\:{x}_{i}\left(t\right)\) exemplifies the place of the \(\:ith\:\)DB at the \(\:tth\) iteration. \(\:\alpha\:\) specifies whether the DB deviates from its early route, with its value defined randomly as 1 and \(\:-1.\) \(\:k\in\:(\text{0,0.2}\)] means constant that denotes the coefficient of deflection, and \(\:b\) refers to constant using the value range of \(\:\left(\text{0,1}\right)\). \(\:{X}^{\omega\:}\) signifies the global poor position, and \(\:\varDelta\:x\) is applied to mimic solar light. If the DB encounters a problem and can no longer continue rolling, it should dance to decide its new rolling path. This behaviour of dancing is outlined as demonstrated:

$$\:{x}_{i}\left(t+1\right)={x}_{i}\left(t\right)+tan\left(\theta\:\right)\left|{x}_{i}\left(t\right)-{x}_{i}\left(t-1\right)\right|$$

(15)

Here, \(\:\theta\:\in\:[0,\:\pi\:]\), and the location is not upgraded if \(\:\theta\:\) captures values like \(\:0,\) \(\:\pi\:/2\), and \(\:\pi\:.\)

Breeding dung beetles

To guarantee a secure background for their offspring, DBs move the dung balls to safe places and hide them before laying their eggs. Equation (16) presents a boundary selection approach to mimic the egg-laying zone of female DBs.

$$\:\begin{array}{c}L{b}^{*}=\text{m}\text{a}\text{x}\left({X}^{*}\times\:\left(1-R\right),\:Lb\right)\\\:U{b}^{*}=\text{m}\text{i}\text{n}\left({X}^{*}\times\:\left(1+R\right),\:Ub\right)\end{array}$$

(16)

On the other hand, \(\:{X}^{*}\) embodies the local best value, and \(\:U{b}^{*}\) and \(\:L{b}^{*}\) characterize the upper and lower limits of the spawning space. \(\:R=1-T/{T}_{\text{m}\text{a}\text{x}}\) and \(\:{T}_{\text{m}\text{a}\text{x}}\) symbolize the maximal iteration counts, and \(\:Ub\) and \(\:Lb\:\)imply the upper and lower limits of the optimizer issue, correspondingly. According to the equation mentioned above, the borders of the egg-laying zone are dynamically established by the \(\:R\) variation. Therefore, the DB’s breeding location is upgraded constantly, as stated by the succeeding mathematical representation:

$$\:{B}_{i}\left(t+1\right)={X}^{*}+{b}_{1}\times\:\left({B}_{i}\left(t\right)-L{b}^{*}\right)+{b}_{2}\times\:\left({B}_{i}\left(t\right)-U{b}^{*}\right)$$

(17)

Here, \(\:{B}_{i}\left(t\right)\) signifies the locality of the \(\:ith\:\)area at the \(\:tth\) iteration. \(\:{b}_{1}\) and \(\:{b}_{2}\) denote two self-governing randomly formed vectors, all using the dimension of \(\:1\)x\(\:D\), whereas \(\:D\) embodies the dimensionality.

Foraging dung beetles

Naturally, while foraging, DBs prefer a safer place in a way equivalent to after they lay their eggs. The unique description of this safer place is re-examined and is characterized by the succeeding Eq. (18):

$$\:\begin{array}{c}L{b}^{b}=\text{m}\text{a}\text{x}\left({X}^{b}\times\:\left(1-R\right),\:Lb\right)\\\:U{b}^{b}=\text{m}\text{i}\text{n}\left({X}^{b}\times\:\left(1+R\right),\:Ub\right)\end{array}$$

(18)

Here, \(\:{X}^{b}\) exemplifies the global finest place, and \(\:U{b}^{b}\)and \(\:L{b}^{b}\) denotes the upper and lower boundaries of the optimum foraging zone, correspondingly. Then, the place of the more minor DB is upgraded as shown:

$$\:{x}_{i}\left(t+1\right)={x}_{i}\left(t\right)+{C}_{1}\times\:\left({x}_{i}\left(t\right)-L{b}^{b}\right)+{C}_{2}\times\:\left({x}_{i}\left(t\right)-U{b}^{b}\right)$$

(19)

\(\:{C}_{1}\) refers to a random variable that follows a standard distribution, and \(\:{C}_{2}\) signifies an arbitrary variable in the interval of \(\:\left(\text{0,1}\right)\).

Stealing dung beetles

This behaviour involves stealing dung balls from another beetle. In the iterative procedure, the location-updated mechanism for the stealing beetle is directed by Eq. (20).

$$\:{x}_{i}\left(t+1\right)={X}^{b}+S\times\:g\times\:\left(\left|{x}_{i}\left(t\right)-{X}^{*}\right|+\left|{x}_{i}\left(t\right)-{X}^{b}\right|\right)$$

(20)

Meanwhile, \(\:S\) signifies a constant, and \(\:g\) denotes a randomly generated vector with dimensions following the standard distributions. Algorithm 1 represents the DBO model.

The DBO method obtains an FF to accomplish enhanced classification performance. It governs a progressive number to epitomize the higher performance of the candidate solutions. The minimization of the classification error rate is measured as the FF, as provided in Eq. (21).

$$\ \begin{aligned} fitness\left( {x_{i} } \right) = & ClassifierErrorRate\left( {x_{i} } \right) \\ = & \frac{{no~of~misclassified~samples}}{{Total~no~of~samples}} \times 100 \\ \end{aligned}$$

(21)



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *