Spontaneous eye blink-based machine learning for tracking clinical fluctuations in Parkinson’s disease

Machine Learning


Subject recruitment and ethical approval

Patients with PD were recruited from the Department of Neurology at Juntendo University School of Medicine between May and December 2021. All patients provided written informed consent after being fully informed of the purpose and procedures of the study. This clinical trial has been approved by the Research Ethics Committee, Faculty of Medicine, Juntendo University (H20-0376) and registered in the University Hospital Medical Information Network- Clinical Trials Registry (UMIN000044246).

The study design and patient population are as follows:

This study is an uncontrolled, open-label, and exploratory clinical study. It consists of 1 week of baseline assessment, habituation to the glasses-type device, and training in symptom diary, followed by a 1 day blink evaluation. On the day of the blink evaluation, the patients were administered a single dose of L-dopa/DCI in the fasting state, following a 12 h discontinuation of anti-PD medication. The blink information, fluctuations in Parkinson’s motor symptoms, and L-dopa pharmacokinetics were evaluated for a maximum of 4 h after administration.

The subjects were 20 patients with advanced-stage PD and fluctuating symptoms, who had been hospitalized for device therapy since they were suffering difficulties in their daily lives due to the wearing-off phenomenon. The authors employed patients who met and did not conflict with the inclusion and exclusion criteria as follows, respectively: Inclusion criteria: (1) Clinically established or probable PD meeting the MDS clinical diagnosis criteria for PD (2015), (2) Stage ≤ III on the Hoehn and Yahr scale in the ON state, (3) Receiving treatment with L-dopa for ≥ 6 months (26 weeks) and showing its effects, (4) Patients with advanced PD who are to be hospitalized for medical evaluation, drug adjustment, and rehabilitation, (5) capable of providing a voluntary written informed consent based upon their sufficient understanding of the research, (6) Patients who are able to join the evaluation without using their own eyeglasses in the case that they regularly use glasses in their daily life. Exclusion criteria: (1) Atypical Parkinsonism syndromes, (2) Dementia or at high risk of it (Mini Mental State Examination (MMSE) score ≤ 20), (3) Contraindicated for concomitant medications (L-dopa/DCI), (4) Hypersensitivity to concomitant medications (L-dopa/DCI) and/or their ingredients, (5) Co-existing psychiatric disease (e.g., depression, bipolar disorder or schizophrenia) and/or clinically significant complications (e.g., cerebrovascular accident, heart disease, chronic respiratory disease, uncontrolled hypertension and diabetes), (6) History of psychiatric disease (e.g., depression, bipolar disorder or schizophrenia) and/or device-aided therapies (i.e., GPi pallidotomy, thalamotomy, and deep brain stimulation), (7) Those who the principal investigator judges to be inappropriate as research subjects

Clinical assessments

The subjects underwent a series of assessments at baseline to obtain data regarding their age, gender, body mass index, disease duration, Mini-Mental State Examination (MMSE) scores, levodopa equivalent daily doses, and stage of the disease according to the Hoehn-Yahr scale, both during and outside of episodes. The following instruments were utilized: the EuroQol 5 Dimensions 5-Level (EQ-5D-5L), the Parkinson’s Disease Questionnaire-39 (PDQ-39), the Movement Disorder Society Unified Parkinson’s Disease Rating Scale (MDS-UPDRS) Parts I, II, III, IV, and the Unified Dyskinesia Rating Scale (UDysRS). To confirm L-dopa levels on the day of blink evaluation, multiple measurements of MDS-UPDRS part III, UDysRS, and bedside score were obtained at the same time as blood collection. The MDS-UPDRS and UDysRS were evaluated by trained experts. Patients were instructed to record bedside scores for their motor symptoms of PD in their diaries, according to a four-point scale: completely off (1), partially off (2), partially on (3), and completely on (4), while discussing the symptoms with their attending physicians.

Pharmacokinetics assessments of levodopa

Following an overnight fast and medication-free period of 12 h, the patients were administered a single dose of L-dopa/DCI, which was the same dose as that for usual regimen. Blood samples were collected for pharmacokinetic analysis of L-dopa at the following time points: before administration and at 15, 30, 45, 60, 90, 120, 150, 180, and 240 min after administration. Plasma L-dopa concentration was determined by the validated method described in Measurement of L-dopa Concentration Section. The maximum L-dopa concentration (Cmax), time to maximum L-dopa concentration (Tmax), and area under the concentration-time curve from time 0 to the last measured time point (AUC0-t) were obtained as parameters of L-dopa pharmacokinetics.

Measurement of L-dopa concentration

The plasma L-dopa concentration was measured by high-performance liquid chromatography (Thermo Scientific™ UltiMate™ 3000 HPLC system). Briefly, 500 μL of plasma was obtained by centrifugation of blood collected in EDTA-2Na tubes (3000 rpm, 10 min, 4 °C), mixed with 50 μL of 60% perchloric acid and centrifuged (12,000 rpm, 40 min, 4 °C). The supernatant was further centrifuged in an ultrafree tube, and the newly obtained purified supernatant (25 µL) was injected into an HPLC system equipped with a WPS-3000 TRS autosampler with a cooling device, an Acclaim™ 120 C18 column (Ф4.6 × 150 mm), and an ECD-3000RS detector. The mobile phase buffer A was prepared by adding 27.6 g sodium phosphate, 680 µL of 0.2 mg/mL nitrilotriacetic acid, and 100 µL tetrahydrofuran to distilled water so that the total volume would be 2 L, and then mixing it with 400 µL 5% SDS, 200 µL ProClin150, and 728 µL phosphoric acid. The mobile phase buffer B was prepared by adding 27.6 g sodium phosphate and 100 µL of 0.2 mg/mL nitrilotriacetic acid to distilled water so that the total volume would be 1 L, and then mixing it with 1060 mL methanol, 4.2 mL 5% SDS, and 6 mL phosphoric acid. The gradient elution was delivered as follows (A:B): 0–2.5 min, 96:4; 2.5–12.5 min, 96:4–62:38; 12.5–21.5 min, 62:38–40:60; 21.5–25.5 min, 40:60–10:90; 25.5–30.5 min, 10:90; 30.5–31.5 min, 10:90–96:4. L-dopa was separated from the buffer solution at a flow rate of 1.0 mL/min and column temperature of 31 °C. Chromatographs were analyzed by Chromeleon 7.2. The limit of detection and quantification were 6.74 pmol/mL and 22.4 pmol/mL, respectively.

Acquisition of pupil data

Pupil data were collected from the eyes of patients using a wearable eye-tracking headset (Fig. 4A) (Pupil Core; Pupil Labs GmbH, Berlin, Germany) for a period of 4 h following L-dopa administration on the day of blink evaluation, with data recorded from 30 min prior to administration until the end of the 4-h period. The headset frame is made of flexible and lightweight plastic (9 g)25. This design did not cause any oppressive or uncomfortable feelings when wearing the device. The headset was equipped with two infrared eye cameras, with one camera dedicated to each eye and a sampling rate of 200 Hz. A USB cable was utilized to establish a connection between the headset and a smartphone (Android OS, version 11) affixed to the patient’s upper arm. The video data of the eyes was recorded with the Pupil Mobile software (version 1.2.3).

Blink data extraction

Pupil and blink data were extracted from recorded videos of eyes with Pupil Player software (v. 3.5.1) (Fig. 4B). Pupil Player detects pupils based on ellipse fitting25 to the infrared video image where a pupil appears as a black ellipse when the eyelid is opened. When the pupil is not detected at all, the value “pupil confidence” based on ellipse fitting was set to 0.0, while the value is 1.0 when the pupil is detected as an ellipse with high accuracy. Pupil confidence for each eye was concatenated by recorded time and a differential filter was applied to extract blink onset and offset. Blinking is detected with a decrease in filtered pupil confidence below the threshold (onset) caused by the concealment of the pupil by the eyelid, and a subsequent increase in filtered pupil confidence above the threshold (offset) caused by eye-opening, occurring at a time less than the filter length. The data were cleaned based on the confidence in the pupil detection to eliminate poor recordings of eye opening as following processes:1. Data Division: We divided the 200 Hz pupil data into 3-min time windows. Each represents an independent data point, focusing on eye blinking characteristics. 2. Criteria for Data Quality:1) If pupil confidence is below 0.2, we consider the eyes closed. 2) If the eyes are closed for > 50% of the time window, or if the average pupil confidence when the eyes are open is below 0.5, we exclude this data point. This is because the patient’s eyes might not be fully open due to reasons like drowsiness or camera setup issues.

Categorization of blink rate changes

PD patients have been reported to show either an increase or decrease in sEBR17,18 in response to L-dopa, we categorized patients into “Increased,” “Decreased,” and “Unchanged” groups based on the changes in blink rate from OFF to ON periods in our observation. In order to categorize these patterns, we first defined ON and OFF periods based on 1-4 bedside scores (BS):

$${\rm{ON}}{\rm{:= }}{\rm{BS}} > 2$$

$${\rm{OFF}}{\rm{:= }}{\rm{BS}}\le 2$$

We then calculated the number of blinks per minute during these periods, denoted as sEBRON for ON periods and sEBROFF for OFF periods. Patients were categorized into “Increased,” “Decreased,” and “Unchanged” groups based on the changes in blink rate from OFF to ON periods. We used a 15% change in sEBRON compared to sEBROFF as an indicator:

$${Increased}:=\frac{{{sEBR}}_{{ON}}}{{{sEBR}}_{{OFF}}} > 1.15$$

$${Decreased}:=\frac{{{sEBR}}_{{ON}}}{{{sEBR}}_{{OFF}}}\le 0.85$$

$${Unchanged}:=1.15\ge \frac{{{sEBR}}_{{ON}}}{{{sEBR}}_{{OFF}}} > 0.85$$

The 15% threshold for defining changes was subjectively and retrospectively chosen to represent our observational results best.

Machine learning procedure and statistical analysis

We aimed to estimate concurrent clinical information from blink features within a specific time window using machine learning techniques (Fig. 4D). In order to achieve this, we employed both the DataRobot platform (Versions: 7.2.8 and 8.0.6; DataRobot, Inc.), which automates model selection and hyperparameter tuning (automated machine learning), and custom Python programs. Ensemble models were intentionally omitted to simplify interpretation.

Data processing

Time-series blink data were segmented into 3 min windows and treated as independent data points. We utilized 1485 cleansed data points from 20 patients for this process. Target variables included: 1.Presence of dyskinesia (binary classification), 2.Bedside score-based ON/OFF state (binary classification), 3.Total MDS-UPDRS Part III score (regression), 4.Plasma L-dopa concentration (regression).

Due to the actual values of target variables being obtained every 15−30 min, we employed linear interpolation for data points within these intervals.

Feature extraction and normalization

We extracted 468 blink-related features from each time window, including base blink features such as blink rate, blink intervals and blink confidence, and their derivatives. The definition of base blink features is summarized in Table 2 and Fig. 4B. The derivatives of base features were extracted by following processes. Also see Fig. 4C and Table 3 for details. First, blink rate and energy—referred to as “base blink features #1”—were obtained as a single value per time window. A baseline correction was performed using the individual’s average before L-dopa administration, which transformed the raw data into three types of features: the original unprocessed value, the difference from baseline, and the absolute difference from baseline. These features were subsequently normalized in two ways: one normalization was applied within each patient’s data to capture intra-individual variability, and another normalization was performed across patients to clarify the relative positioning. For the training phase, the normalization utilized data from all patients included in the training, and for the testing phase, the test patient’s features were normalized using data from all patients. Similarly, base blink features #2—comprising interval, duration, confidence, and depth—were calculated for each blink within the time window, allowing the derivation of statistical parameters such as mean, median, standard deviation, maximum, and minimum. In addition, the mean and standard deviation from all blinks observed during the recording for each individual were used to categorize these features into five groups (LOW, MID-LOW, MID, MID-HIGH, and HIGH), according to rules detailed in Table 3. The frequency of each category was then counted within each time window; for instance, if three LOW and two MID blinks were observed for blink confidence, the features were recorded as CONFIDENCE_FRQ_LOW = 3 and CONFIDENCE_FRQ_MID = 2. Furthermore, the relative frequency and the ratio of each category to the most frequently observed category (MID) were computed. In the example above, CONFIDENCE_FRQ_LOW_REL = CONFIDENCE_FRQ_LOW / CONFIDENCE_FRQ_MID = 3 / 2 = 1.5. Similar baseline correction and normalization procedures as for base blink features #1 were applied, resulting in a total of 468 blink-related features.

Table 2 Summary of base blink features, background and reference features
Table 3 Construction of derivative features. Also see the text and Fig. 1C for a detailed feature engineering procedure

In addition to blink features, non-blink features such as plasma L-dopa concentration, elapsed time after L-dopa administration, and patient age were included. Plasma L-dopa concentration is an invasive but reflective indicator of Parkinson’s disease (PD) motor symptoms26, while elapsed time after L-dopa administration is easily obtained and anticipated to correlate with PD motor symptoms. Patient age was also included, as it is known to relate to both blink dynamics and PD symptoms23,27. To prevent overfitting due to predictable data collection intervals and small sample size, Gaussian noise (σ = 30) was added to elapsed time, while Gaussian noise (σ = 3) was added to patient age. These features were normalized in two ways: first, within-patient normalization was applied to capture each patient’s relative variability; second, across-patient normalization was performed to clarify relative positioning across the patient group (Fig. 4B, C; Tables 2 and 3).

Model Construction and Evaluation

We performed model construction and evaluation using a leave-one patient/group-out approach. For ON/OFF state, MDS-UPDRS Part III score and plasma L-dopa concentration, models were trained using data from 19 patients and evaluated on the remaining patient. For dyskinesia, since only seven patients exhibited this symptom, a leave-one-patient-out evaluation with AUCROC was not feasible. In order to address this, we assigned one or two non-dyskinetic patients to each dyskinetic patient, forming two- or three-patient groups for leave-one-group-out evaluation. Non-dyskinetic patients were selected from different blink rate change patterns (i.e., Increased, Decreased, or Unchanged) randomly to ensure diversity. We prepared 14 groups to prevent bias, ensuring that each patient was included in two groups without being paired with the same patient.

During each training phase of leave-one patient/group-out approach, 20 blink-related features were selected for each target variable by a two-step approach. First, the Boruta algorithm was used to exclude features with contributions that were statistically lower than random noise28. Next, SHAP (SHapley Additive exPlanations) analysis29 was applied to iteratively remove features with the lowest contributions at a 10% exclusion rate until 20 features remained.

Evaluation Metrics and Statistical Analysis

The performance of models was evaluated based on the task. For binary classification tasks, including ON/OFF state and dyskinesia, models were evaluated based on the LogLoss metric, and the area under the receiver operating characteristic (ROC) curve (AUCROC) was calculated for each eligible patient or patient group and compared. For dyskinesia, ROC curves were generated for both ON state and the entire evaluation period for 14 groups. For ON/OFF state, ROC curves were generated for 19 patients. One patient who was entirely in the OFF state and did not show any ON state was omitted. For regression tasks, including MDS-UPDRS Part III scores and plasma L-dopa concentration, performance was assessed using the Root Mean Squared Error (RMSE). Then Spearman’s correlation (ρ) between predictions and actual values was calculated for each patient and compared, with the proportion of patients with positive (ρ > 0) and statistically significant ρ (test of no correlation (two-sided, p < 0.05)). In order to examine the extent to which the predicted data could replicate original trends when restored to a time series, we calculated Derivative Dynamic Time Warping (DDTW) for each patient30.

Predictive performance was assessed across three different feature combinations:1. Selected blink features (B), 2. Blink features + background information (elapsed time post-L-dopa and patient age) (B + BG), 3. Plasma L-dopa concentration alone (L). Additionally, we evaluated the predictive performance after post-processing raw predictions with a 15 min moving average for practical application (smoothed). Results were compared using paired t-tests with Bonferroni’s correction. All possible feature combinations were tested, i.e., B, B + BG, and L, both with and without smoothing.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *