Study subjects
The study protocol received approval from the institutional ethical committee of SRM Medical College Hospital and Research Centre (834/IEC/2015). The written informed consent and detailed questionnaire were obtained from all the enrolled participants (N = 200) to examine their health status before recruiting for the clinical study. After the strict scrutiny of questionnaires, 40 participants with confounding factors such as pregnant or nursing women, cardiovascular problems, renal failures, fever, thyroid disorders, and anemia were excluded from the clinical study. The remaining 160 participants recruited for the clinical study. The blood sample is collected for all the recruited participants in fasting and postprandial (2 h) conditions to measure their glucose profile. We acquired the tongue thermal and visible images from 160 recruited participants. Based on the diabetes diagnostic criteria by ADA15, the study subjects (N = 160) are categorized into two groups, namely.
-
(1)
Group I: Normal (N = 80), comprised of age and sex-matched subjects with a Male: Female ratio of 1:2, and a mean age ± standard deviation of 41.23 ± 10.82 years.
-
(2)
Group II: Type II DM (N = 80), with a Male: Female ratio of 1:2 and a mean age ± standard deviation of 42.95 ± 9.63 years.
The proposed framework for a study focused on pre-screening for Type II diabetes mellitus is as follows (Fig. 1).

The envisaged study design for pre-screening diabetes mellitus.
Physiological and biochemical measurements
Anthropometrical variables, including body height (cm), body weight (cm), body mass index (BMI, kg m−2), hip and waist circumferences (cm), systolic blood pressure (SBP, mmHg), and diastolic blood pressure (DBP, mmHg), were assessed for all participants. The FBG, PPBG, and HbA1c are the standard biochemical tests performed using participants’ extracted blood samples.
Thermal and visible tongue images acquisition and analysis
From the recruited participants (N = 160), the thermal and visible tongue images are captured using a thermal infrared camera (FLIR A305 SC) and digital single-lens reflex (DSLR) camera (Nikon D5300, 24.2 megapixels), respectively. The FLIR A305SC typically has a thermal resolution of 320 × 240 pixels. It has a thermal sensitivity, typically around 0.05 °C. This means it can detect even small temperature differences accurately. The FLIR A305SC cameras come equipped with an integrated IR lens featuring an 18 mm focal length. The field of view (FOV) is defined as 25° × 18.8°, with an Instantaneous Field of View (IFOV) of 1.36 mrad. These cameras exhibit an accuracy level of ± 2 °C or ± 2% of the reading. The camera was configured with a temperature range of 22–45 °C. We maintained the ambient room temperature between 22 and 23 °C with a relative humidity of 50%. The emissivity of the tongue, when using an infrared thermal camera, is often set to a standard value of around 0.98, corresponding to the emissivity of human skin. The participants are requested to be seated for 15 min to equilibrate themselves with the ambient room temperature. The thermal and visible images are acquired during the fasting condition. The distance between the camera and the subject’s tongue is consistently maintained at 0.3 m34. Before initiating the imaging process, subjects are instructed to open their mouths widely and extend their tongues downward for 1 min. To prevent artifacts and enhance the background, a black cloth is positioned in front of the patient’s mouth. The dimensions of the background were set according to the thermal camera’s field of view and the distance between the mouth and the camera. The stabilization time of about 2 min is required between the camera and the patient during the thermal measurements35. Skin temperature can be influenced by various factors, including prolonged exposure to specific environmental conditions such as extreme temperatures, humidity, and solar radiation, engaging in extended periods of physical activity36. The evaporation of moisture, such as sweat, on the skin’s surface can have a significant impact on temperature measurements obtained by a thermal camera. The elevated temperatures be attributed to the reduced evaporation rate, possibly caused by inadequate saliva secretion in tongue of diabetic subject. The thermal tongue images are analysed using FLIR Version 2.0 and MATLAB version R2021a, (Math Works, California, USA) with deep learning package ResNet 50, VGG16. We have chosen the rainbow palette with a constant temperature scale of 28.5–36.9 °C for all the tongue thermograms. According to TCM, the middle portion of the human tongue is connected to the human stomach and Pancreas, as it involves digestion and diabetic conditions37. But the link between the middle part of the tongue and the pancreas is significant because it’s associated with taste receptors that can help detect sweet flavors, which in turn may influence the release of insulin from the pancreas to regulate blood sugar levels. The upper part of the tongue is associate with kidney. The tip of the tongue is associated with the heart and lungs in TCM and certain alternative health practices. The left and right-side associated regions are related to liver and gall bladder. So, the region of interest (ROI) is positioned at the central part of the tongue (Fig. S1). A square area tool of dimension 50 × 50 pixels is used for analysing the fused tongue images. A square ROI of uniform size of 50 × 50 pixels is fixed semi-automatically on the central region of tongue in visible, thermal and fused tongue images. The ROI is cropped manually and features are extracted using MATLAB programming. The experimental set up illustrates the thermal image acquisition process (Fig. 2).

Experimental set up of thermal image acquisition.
Tongue image fusion: DWT method
The fusion of human thermal (FLIR) and visible tongue (DSLR) images are performed using discrete wavelet transform based on fusion rules as follows:
First, the thermal (IR) and visible tongue (DSLR) images need to be pre-processed to ensure they are aligned and have the same dimensions. The thermal image has a real size of 320 × 240 pixels, while the DSLR image has a real size of 390 × 280 pixels. The camera time is not synchronized, with the thermal image captured first and a 2-s delay before acquiring the visible image. Despite potential movement of the tongue between the thermal and visible images, geometric transformations such as translational shifting and rotation (up to 5 degrees) are applied to align the images during the fusion process.
Discrete wavelet transform (DWT) is a mathematical technique employed to decompose an image into various scales and orientations, capturing both high-frequency and low-frequency components. The discrete wavelet transform (DWT) fusion method is often chosen over other fusion techniques due to its ability to efficiently integrate information from multiple sources while preserving important features. DWT allows for multiresolution analysis, which means it can decompose images into different frequency components at varying levels of detail. It enables the representation of both coarse and fine features, capturing a wide range of information. DWT provides a sparse representation of images, which means it concentrates most of the signal energy in a few coefficients. This sparsity property is beneficial for fusion because it facilitates the extraction of relevant information while reducing redundancy. It can localize features in both space and frequency domains. This capability is essential for fusion tasks as it helps preserve the spatial and spectral characteristics of the input data. DWT offers computational efficiency compared to other fusion techniques, especially when dealing with large datasets or real-time processing requirements.
Despite its advantages, the DWT fusion method also presents some limitations and challenges. The DWT has limited directionality because it relies on a predefined set of wavelet basis functions, which may not always capture the directional information present in the input data effectively. The discrete nature of the DWT introduces boundary effects, where discontinuities at the edges of images can lead to artifacts in the fused result. DWT fusion results can be sensitive to scale and shift variations in the input data, particularly when dealing with images acquired under different conditions or sensor configurations. The multiresolution nature of DWT involves a trade-off between resolution and information loss, where higher levels of decomposition offer better frequency resolution but may lead to increased loss of spatial or spectral details.
Apply DWT separately to the thermal and visible tongue images to obtain their respective wavelet coefficients at different scales and orientations. Fusion rules are used to combine the wavelet coefficients from both images. Example if the rule is Max-Max, select the maximum value of the corresponding coefficients from both images at each scale and orientation. The selection rule is based on choosing the coefficients from one modality based on certain criteria, for example, selecting thermal coefficients for low-frequency components and visible coefficients for high-frequency components. After applying the fusion rule to the wavelet coefficients, perform the inverse DWT to obtain the fused image. This fused image will contain combined information from both the thermal and visible tongue images. The wavelet filter used for tongue image fusion is Daubechies (dB) filter with order two, and the decomposition level is 2. The detailed illustration of tongue image fusion are as follows: (Fig. 3).

Tongue image fusion using discrete wavelet transform (DWT) method.
The fusion of visible and thermal tongue images is performed based on the fusion rules (Table S1). There are nine different fusion rules for fusing visible and thermal tongue images. For illustration purpose, the proposed study elaborates the mean–max fusion rule. According to the mean–max rule, mean is considered as approximate co-efficient, which contains low frequency information. Max is called detailed co-efficient, containing high frequency level information. The low frequency information generally suppresses the average noise present in the image based on adopting the simple average method. A high frequency component extracts the detailed information related to curves, lines, and contours present in the source image. Hence in the proposed study, the low frequency component from the visible tongue image is fused with the high frequency component from the thermal tongue image. In the end, the reconstructed fused tongue image is acquired through the application of the inverse wavelet transform. The Eq. (1) represents the mean-max fusion rule as given below as follows:
$${\text{I}}_{{{\text{fuse}}}} \left( {{\text{x}},{\text{y}}} \right) = {\text{W}}^{ – 1} \left[ {\upphi _{{\left( {{\text{IRT}},{\text{VIS}}} \right)}} \left\{ {{\text{W}}_{{\text{T}}} \left( {{\text{I}}_{{{\text{IRT}}}} \left( {{\text{x}},{\text{y}}} \right)} \right.,{\text{W}}_{{\text{T}}} \left( {{\text{I}}_{{{\text{VIS}}}} \left( {{\text{x}},{\text{y}}} \right)} \right.} \right\}} \right]$$
(1)
whereas Ifuse(x,y)—fused tongue output image, W−1—inverse discrete wavelet transform, ϕ—fusion rule mean–max, WT—wavelet transform, IIRT(x,y)—thermal tongue input image, IVIS(x,y)—visible tongue input image.
$$\upphi _{{\left( {{\text{IRT}},{\text{VIS}}} \right)}} = {\text{mean}},\max$$
$$\upphi _{{\left( {{\text{IRT}},{\text{VIS}}} \right)}} = {{\left( {{\text{IRT}} + {\text{VIS}}} \right)} \mathord{\left/ {\vphantom {{\left( {{\text{IRT}} + {\text{VIS}}} \right)} 2}} \right. \kern-0pt} 2}\;\,{\text{if IRT}} \ne {\text{VIS}}$$
(1a)
$$\upphi _{{\left( {{\text{IRT}},{\text{VIS}}} \right)}} = {\text{IRT}}\;\,{\text{if IRT}} = {\text{VIS}}$$
(1b)
The mean-max rule calculates the mean (average) of the two values if IRT and VIS are not equal. It takes the maximum value if IRT and VIS are equal, essentially preserving the maximum value.
According to mean-max fusion rule, mean is derived from approximate coefficient and max is obtained from detailed coefficient
$${\text{I}}_{{{\text{fuse}}}} \left( {{\text{x}},{\text{y}}} \right) = {\text{W}}^{ – 1} \left[ {{\text{mean}},\max \left\{ {{\text{W}}_{{\text{T}}} \cdot \left( {{\text{Image}}_{1} \left( {{\text{x}},{\text{y}}} \right)} \right.,{\text{W}}_{{\text{T}}} \cdot \left( {{\text{Image}}_{2} \left( {{\text{x}},{\text{y}}} \right)} \right.} \right\}} \right]$$
(2)
Statistical feature extraction
The thermal, visible, and fused tongue images were converted into grayscale images to extract the statistical features. The Gray level co-occurrence matrix (GLCM) relies on a statistical technique employed to examine texture features, providing insights into the spatial pixel relationships38,39,40,41. We extracted the statistical parameters such as mean, contrast, standard deviation, correlation, energy, entropy, homogeneity, skewness, variance, and kurtosis from the thermal, visible, and fused tongue images using the GLCM algorithm. The extracted statistical features from the thermal, visible, and fused tongue images were provided (Table S2).
Machine and deep learning algorithms
Machine learning classifiers
The SVM, linear discriminant analysis (LDA), k-nearest neighbour (k-NN) and Visual Geometry Group Net (VGG16) and ResNet50 were used to perform the classification of diabetes from the fused tongue image. The SVM classifier performs both linear and non-linear (using kernel function) classification. It uses hyperplanes to define the boundaries that exhibit better classification accuracy for minimum datasets. SVM is a binary classifier that maximizes the margin to determine the hyperplane that separates the two classes42. LDA is a dimensionality reduction method that separates two or more groups and extends the features in the higher dimensional space to the lower dimension space. It measures various linear features within and between class-scatter matrices43. k-NN is used for classification and regression problems. It works on the concept that nearer objects are mainly expected to be surrounded by a similar category44. In machine learning classifiers, total data used is 160 (80 for diabetic and 80 for Normal). The data split used for training is 70% (112 images), validation is 15% (24 images) and testing is 15% (24 images).
Convolution neural network
Convolutional neural network (CNN) is a type of deep neural network that demonstrates exceptional performance in medical image classification by extracting and learning complex high-level features from the images45. The overall architecture of the VGG16 model for Tongue image classification are explained as follows (Fig. 4). VGG16 comprises 13 convolutional layers arranged into five convolutional blocks. Each block consists of multiple 3 × 3 convolutional layers, followed by a max-pooling layer with a pool size of 2 × 2. The convolutional layers are structured to capture various levels of image features, progressively learning more intricate patterns. After the convolutional blocks, VGG16 has three fully connected layers. Each fully connected layer is followed by a rectified linear unit (ReLU) activation function. Before the fully connected layers, there is a flattening layer that converts the 3D feature maps into a 1D vector. The final layer is a SoftMax activation layer, which produces the probability distribution over the different classes. The number of neurons in the output layer aligns with the number of classes in the classification task. The exclusive use of 3 × 3 convolutional filters across the network facilitates a deeper architecture, and the repeated stacking of convolutional layers aids in learning hierarchical features. The stochastic gradient descent (SGD) optimization algorithm is used in VGG16 Net to update the neural network weights with a 0.01 learning rate46. Categorical Cross Entropy is employed as the loss function in VGG16.

Pre-trained VGG-16 Net for tongue image classification.
ResNet50 is a deep CNN architecture that has garnered significant popularity in the field of computer vision, particularly for tasks such as image classification and object detection47. The key innovation of ResNet50 lies in the incorporation of residual blocks, aiming to overcome the vanishing gradient problem and facilitate the training of extremely deep neural networks. Residual blocks, the core building blocks of ResNet50, include skip connections, also known as shortcut connections, which bypass one or more convolutional layers. It consists of 48 convolutional layers and 1 fully connected layer, making it a very deep architecture. The input layer accepts input images with a standard size, typically 224 × 224 pixels. The initial layers of ResNet50 comprise standard convolutional layers and pooling layers, designed to extract fundamental features from the input image. The core building blocks of ResNet-50 are residual blocks. Each block contains two or three convolutional layers along with skip connections. Multiple residual blocks are stacked together to form the deeper layers of the network. ResNet-50 includes four stages, each with a different number of residual blocks. After the convolutional layers, a global average pooling layer is used to reduce the spatial dimensions of the feature maps. A fully connected layer processes the output of the global average pooling layer to generate the final classification scores. The SoftMax activation function is then applied to the final layer to transform the raw scores into class probabilities. The output layer provides the final predictions for the input image, indicating the probabilities of different classes. Stochastic gradient descent (SGD) optimizer is used as a hyperparameter in ResNet50 to update the parameters using 0.01 as learning rate. Categorical Cross Entropy is utilized as the loss function in ResNet50.
Google Cloud Platform, commonly known as GCP, is a comprehensive suite of cloud computing services offered by Google. It encompasses a broad array of tools and services for computing, storage, data analytics, machine learning, and more. Google cloud platform is used for training and validation of the CNN model. Dedicated hardware is required for training the neural network weights of the VGG16 net and ResNet50. Hence NVIDIA TESLA V100 based cloud GPU with 16 GB RAM is initiated for the accelerated learning and training process. The CNN models for diabetes detection make use of TensorFlow’s image data generator API, specifically version V2.13.0, for efficient execution. TensorFlow, often combined with scikit-learn for machine learning tasks, includes the Keras API for deep learning. OpenCV is commonly used for image processing and computer vision tasks. The fused tongue thermograms are resized to size of 224 × 224 with 3 channels (RGB). The data augmentation techniques such as translational shifting, horizontal flipping, shearing, zooming, and rotational were used to increase the number of inputs thermograms for training deep learning models. As there are five augmentation techniques used, totally 960 images are used after augmentation (160 images (both Normal and diabetes) × 5 = 800 augmented images + 160 original images = 960). Hence the ratio of augmented to original image is 1:6.
Data split
The thermogram datasets are divided into three disjoint sets namely training 672 images (70%) used for training, 144 images (15%) used for validation and 144 images (15%) dataset for testing. The same data split is used for both VGG16 model and ResNet50 model. The model is trained for 10 epochs and during each epoch cycle, the CNN will be trained with the train data and gets checked with validation data to get error. Based on this error, the network weights will be varied and retrained for next epoch. Other than weights, other parameters will not be tuned. Early stopping criteria is used to stop the training. The Validation accuracy is used as the performance metric in early stopping with patience = 2. The training process halts when the selected performance measure no longer shows enhancement. Finally, the fully trained model is analysed for the classification performance with the testing dataset.
Performance evaluation
For the evaluation of image fusion performance, the image quality metrics such as normalized cross-correlation (NCC), mean square error (MSE), peak signal–noise ratio (PSNR), normalized absolute error (NAE), average difference (AD), maximum difference (MD), signal to noise ratio (SNR), structural content (SC) and structural similarity index (SSIM) were used in the proposed study. The definition of image quality metrics48 briefed as follows:
Mean Square Error (MSE) quantifies the error of an image by assessing the disparity between the original input image and the processed output image. The lesser value of MSE denotes better performance.
$$MSE = \frac{1}{mn}\mathop \sum \limits_{k = 1}^{m} \mathop \sum \limits_{k = 1}^{n} \left( {X_{k,l} – Y_{k,l} } \right)^{2}$$
(3)
Peak signal-to-noise ratio (PSNR) is the ratio of the maximum power (peak value) of the information in an image to the power of noise in the image. A higher PSNR value indicates less noise in the image and signifies higher quality in the processed image.
$$PSNR = 10\log_{10} \frac{{p^{2} }}{MSE}$$
(4)
Normalized cross-correlation measures the degree of similarity between the processed and original images.
$$NCC = \mathop \sum \limits_{k = 1}^{m} \mathop \sum \limits_{l = 1}^{n} \frac{{\left( {X_{k,l} *Y_{k,l} } \right)}}{{X_{k,l}^{2} }}$$
(5)
Normalized absolute error is similar to MSE, but they have subtle differences in the values.
$$NAE = \mathop \sum \limits_{k = 1}^{m} \mathop \sum \limits_{l = 1}^{n} \frac{{\left| {X_{k,l} – Y_{k,l} } \right|}}{{\left( {X_{k,l} } \right)}}$$
(6)
Average difference estimates the mean differences between the processed and original images. The value of AD should be as less as possible, and the ideal value is 0.
$$AD = 1/mn\mathop \sum \limits_{k = 1}^{m} \mathop \sum \limits_{k = 1}^{n} \left[ {X\left( {k,l} \right) – Y\left( {k,l} \right)} \right]$$
(7)
The maximum difference measures the pixel-wise maximum differences between the original and processed output images.
$$MD = Max\left( {\left| {X_{k,l} – Y_{k,l} } \right|} \right)$$
(8)
Structural content defines the proportion of the sum of the squares of the reference input and processed output image. Let image with the size of m × n matrix
$$SC = \frac{{\mathop \sum \nolimits_{k = 1}^{m} \mathop \sum \nolimits_{l = 1}^{n} \left( {X_{k,l} } \right)^{2} }}{{\mathop \sum \nolimits_{k = 1}^{m} \mathop \sum \nolimits_{l = 1}^{n} \left( {Y_{k,l} } \right)^{2} }}$$
(9)
whereas m—number of columns, n—number of rows, X—fused or processed image, Y—thermal/visible tongue image, k, l—pixel row and column index.
The structural similarity index (SSIM) is utilized to measure the similarity between two images. It assesses the perceived quality of an image by comparing its structural information to a reference image.
$$SSIM = \frac{{\left( {2\mu_{k} \mu_{l} + C_{1} } \right)\left( {2\sigma_{k} l + C_{2} } \right)}}{{\left( {\mu_{k}^{2} + \mu_{l}^{2} + C_{1} } \right)\left( {\sigma_{k}^{2} + \sigma_{l}^{2} + C_{2} } \right)}}$$
(10)
where C1 and C2 are smoothing constant or regularization parameters
The baseline parameters, anthropometrical variables, tongue temperature, and extracted features from thermal, visible, and fused tongue images are provided as the input to SVM, k-NN, and LDA to perform the classification task. The thermal, visible, and fused tongue images were provided directly as input variables to the CNN to perform the classification task. The area under the curve (AUC) was derived from the receiver operating characteristic (ROC) curve of the classifier. The classifiers performance was evaluated by the assessment metrics such as sensitivity, specificity, accuracy, positive predictive value (PPV), and negative predictive value (NPV).
$${\text{Sensitivity}} = \frac{TP}{{TP + FN}} \cdot 100$$
(11)
$${\text{Specificity}} = \frac{TN}{{TN + FP}} \cdot 100$$
(12)
$${\text{Accuracy}} = \frac{TP + TN}{{TP + FP + TN + FN}} \cdot 100$$
(13)
$${\text{PPV}} = \frac{TP}{{TP + FP}} \cdot 100$$
(14)
$${\text{NPV}} = \frac{TN}{{TN + FN}} \cdot 100$$
(15)
Intra- and inter-observer variability
To study the reproducibility and reliability of the imaging process in Tongue thermogram, the intra-observer variability and inter-observer variability was performed using Bland–Altman plot based on temperature measurements from the Tongue thermogram.
In Inter-observer variability, the temperature measurement was made by two different observers using Tongue thermogram. In intra-observer variability, the temperature measurement was made by same observer at different timings to validate the reproducibility.
Statistical analysis
The data were presented as the mean ± standard deviation (SD). The Shapiro–Wilk test was conducted to assess data normality. To identify significant differences among the groups, the Student’s t-test was employed. The data analysis was carried out using SPSS version 21.0 software, Chicago, USA.
Ethical approval
All procedures performed in studies involving human participants by the ethical standards of the institutional research committee and with the 1964 Helsinki declaration and its later amendments or comparable ethical standards.
Informed consent
Informed consent was obtained from all participants included in the study.
