New AI-driven studies reveal machine learning models that have progressed not only to identify known Alzheimer's disease genes, but also to find six new risk variants.
Research: Machine learning in the genetics of Alzheimer's disease. Image credit: Kateryna Kon/Shutterstock.com
Statistical tools are essential to unlocking the genetic foundations of complex medical conditions. There hasn't been much progress beyond the linear additive model. However, recent papers have been published Natural Communication We discuss the results of applying machine learning (ML) to genomic data from a large cohort of patients with Alzheimer's disease (AD) in Europe.
introduction
Genome-Wide Association Studies (GWAS) have pioneered deeper insights into genetic variation as risk factors for AD. These variants are considered for a polygenic risk score (PRS) that helps to predict disease risk.
These tools are designed with the assumption that variants predict results uniformly. Whether these variants occur at the same or other gene locus, the risk associated with the individual variants is added. This ignores the knowledge that risk can be altered by interactions between variants and with other risk factors.
For example, advertising research shows that it is different. apoe Variants alter disease characteristics and the type of immune cell response to abnormal neuronal proteins. Genetic studies show differences apoe Expression leads to associations between different AD genes and changes in age at diagnosis.
As GWAS sample size increases and PRS plateau forces increase, new platforms applying advanced computational resources are essential to narrow down the greatest benefits from the large currently available data and to better examine the genetic foundations of AD. Artificial intelligence in ML models has been applied in several studies. However, the small sample size has significantly increased the risk of bias.
The current study attempted to address this using the largest genome-wide dataset currently available.
About the research
In this study, researchers trained three types of models. It is well known in the field and is highly efficient.
- Gradient boost machine (GBMS)
- Biological pathway-based neural networks (NNS)
- Model-based multifactor dimension reduction (MB-MDR).
The aim was to assess the effectiveness of each algorithm in performing three types of tasks.
- Reproduction of previous survey results
- Find new disease-related loci overlooked by GWAS
- High-risk individual predictions
This study used rigorous cross-validation, splitting of multiple random train tests, and careful adjustments for confounding factors such as gender, age, genotyping centres, and demographic structure.
result
Replica of previous findings
Regarding initial objectives, the findings showed that ML captured all genetic variants across the genome of the training set. Furthermore, although the sample size was only in the 20th century, we identified 22% of AD-related variants reported in a larger GWAS meta-analysis. Therefore, this study sets a benchmark for ML-based genome-wide methods.
The ability of ML models to replicate findings from much larger GWAS highlights that flexible models can recover a significant portion of known genetic risk in a small sample.
Identifying gene loci
Secondly, ML was correctly identified apoe As a risk factor for AD. It correctly captured lead single nucleotide polymorphisms (SNPs) causally associated with AD. Beyond the methods, ML highlighted the read SNPs of multiple important genes in AD. MB-MDR 1 D found 20 very stable SNPs. apoe Areas with all possible train test divisions.
The model also identified six new loci replicated in an unrelated data set. These loci encode genes such as ARHGAP25, ly6hand Cog7. GBM has identified most new loci.
A new association has been detected AP4E1close to someone already known sppl2a Trajectory. AP4E1 It encodes part of the protein key for amyloid metabolism, and its deficiency may promote beta-amyloid formation and increase the risk of AD. The neural network approach also highlighted additional novel trajectories (SOD1) Possible biological links with AD pathology.
Ad Status Prediction
All models predicted AD status with comparable accuracy. GBM was most strongly correlated with NN and MDRC 1d. Although it was weakly correlated with NNS, PRS was strongly linked to GBMS.
GBM and PRS were excellent at predicting cases that differ from controls. Predictions were validated using random training and test data relocation to show high reproducibility.
Women were overrepresented among predicted cases, as expected from the majority of women in the data. GBM was the exception, with similar proportions of males and females in both cases and controls.
All models' predictions remain stable across different cohorts and repeated random splits, suggesting that the findings are not driven by overfitting or technical artifacts.
Comparison with GWAS
Investigators compared the major ML detected variants with all significant AD-related SNPs reported in the meta-analysis. Of the 130 previously reported genes corresponding to 86 loci, one or more ML algorithms picked up 19. All models have been identified apoetwo models detected seven loci.
Leave apoe Regions from the training dataset led to the identification of more known AD risk genes, but were less accurate. If only the current data was used, one or more ML models identified each GWAS detected SNP in the training dataset.
Higher priority ML-identified SNPs were more enriched in microglia and astrocyte regions. These were involved in a variety of AD-related pathways, including regulation of AD Hallmark beta amyloid proteins or alterations in protein concentrations such as LY6H. This molecule binds to acetylcholine receptors involved in neurotransmission, and its level in cerebrospinal fluid correlates with the severity of AD. Others are tracked for glycosylation abnormalities involved in Ad Tau protein processing.
The way ML models rank the importance of SNPs (e.g., SHAP values for GBM, permutated P values for MB-MDR, or network weights for NNs) does not always translate directly to traditional GWAS significance, reflecting the fundamental differences in feature selection between ML and traditional statistics.
The importance of research
Stylish studies that exert this ability highlight that given the large dataset available to ML, it is possible to predict AD-related genetic variations on par with traditional genome-wide methods.
The moderate prediction accuracy of GWAS meta-analysis may be due to heterogeneity in the included studies, reflecting differences in multiple related characteristics. A more uniform sample provides a higher odds ratio than a clinical sample. Some SNPs identified by the ML model may have detectable effects only under specific cohorts or specific conditions, which may not be visible in large, heterogeneous external data sets.
This also explains why all SNPs identified by the ML model were unable to replicate in the external dataset. Their effects can only be significant in certain circumstances, and cannot demonstrate the importance of the entire genome in very different studies with different circumstances.
Nevertheless, the novel SNPs here influenced biologically plausible pathways. Further research is essential to understand how to identify important SNPs from those captured in a variety of ways.
Conclusion
“Our results show that machine learning methods can achieve predictive performance comparable to classical approaches in genetic epidemiology.. In addition to predicting risk, they identified new loci that they missed with the reproducible approach used here.
Overall, this study illustrates the promises and current limitations of ML in AD genetics. It offers a valuable addition to GWAS, but also highlights the need for careful interpretation, replication and even methodological refinement.
The current study opens ways to develop and validate future ML models and complements traditional methods in AD genetic research.
Download the PDF copy now!
