GLARE: Uncovering Hidden Patterns in Spaceflight Transcriptomes Using Representation Learning

Machine Learning


GLARE: Uncovering Hidden Patterns in Spaceflight Transcriptomes Using Representation Learning

Overall pipeline of GLARE: Gene LAb Representation Learning Pipeline. (a) Diagram of GLARE. It starts with a validation study and uses k-means clustering to detect outliers and perform preprocessing. With a clean dataset, GLARE offers options for representation learning, from PCA to state-of-the-art SAE pre-trained on high-throughput single-cell data. The obtained data representations are then processed by ensemble clustering to find hidden patterns in the data. The results of the validation study and ensemble clustering are used for post-pipeline analysis. (b) Diagram of the model architecture of SAE used in training both with and without pre-training. (c) Ensemble clustering with three basic clustering algorithms based on different statistical methods. Evidence accumulation clustering is used to derive consensus clusters from these algorithms. — biorxiv.org

Spaceflight studies offer new insights into biological processes through exposure to stressors outside the evolutionary pathway of terrestrial organisms. Despite limited access to the space environment, numerous transcriptomics datasets from spaceflight experiments are now available. NASA Genetics Research Institute The data repository will provide public access to these datasets, facilitating further analysis.

Although various computational pipelines and methods have been used to process these transcriptomic datasets, learning model-driven analysis has not yet been applied to the broad scope of such spaceflight-related datasets.

In this study, we propose an open-source framework, GLARE: GeneLAb Representation Learning Pipeline, which consists of training different representation learning approaches, from manifold learning to self-supervised learning, to improve performance on downstream analytical tasks such as pattern recognition. To demonstrate the utility of GLARE, we apply it to gene-level transcriptome values ​​obtained from results of the CARA spaceflight experiment, an Arabidopsis root tip transcriptome dataset spanning light, dark, and microgravity treatments.

GLARE not only validated the results of the original study on cell wall remodeling, but also revealed additional patterns of gene expression affected by the treatment, including evidence of hypoxia. This study suggests that there is great potential to complement insights gained from earlier studies of spaceflight omics-level data with further analyses powered by machine learning.

Analysis of hypoxia clusters found in FLT clustering results. (a) Heatmap of normalized FPKM values ​​on hypoxia clusters. (b) Enriched ontology of hypoxia clusters from Metascape. (c) Stress Knowledge Map (SKM) for five transcription factors (TFs) in the hypoxia cluster: “DREB2A”, “RHL41/ZAT12”, “MYC2”, “RRTF1/ERF109”, “STZ/ZAT10”. — biorxiv.org

GLARE: Discovering Hidden Patterns in Spaceflight Transcriptomes Using Representation Learning, biorxiv.org

Astrobiology, Genome Science,

Explorers Club Fellow, former NASA Space Station Payload Manager/Astrobiologist, Away Team, Journalist, former mountain climber, Synesthete, mixed Na'Vi-Jedi-Freman-Buddhist, ASL, Devon Island and Everest Base Camp veteran, (male) 🖖🏻



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *