Amyotrophic lateral sclerosis (ALS) is a progressive neurodegenerative disorder of upper and lower motor neurons leading to weakness, atrophy, and respiratory failure. Early, minimally invasive diagnostic biomarkers remain lacking. Dysregulation of post-transcriptional modifications, particularly m6A (N6-methyladenosine) on mRNA, can affect stability, splicing, and translation and has been implicated in other neurodegenerative diseases. Single-nucleotide polymorphisms that influence RNA modification status (m6A-SNPs) may alter gene expression and represent candidate molecular markers for ALS.
The study used a multi-omics integration strategy combining large-scale blood cis-eQTL data, curated m6A-SNP annotations, and bulk whole-blood transcriptomic datasets from ALS patients and controls. A machine-learning pipeline—random forest followed by LASSO logistic regression—was applied to identify robust diagnostic marker genes from candidate m6A-SNP–related genes. The derived diagnostic model was trained in one cohort and externally validated in an independent cohort without coefficient re-estimation. Additional analyses included immune-cell deconvolution, in silico prediction of m6A sites, annotation of RNA-binding protein (RBP) interactions near SNP loci, and exploratory clinical validation of gene expression and global m6A levels in peripheral blood samples.
Key data inputs:
Platform-specific preprocessing performed by original depositors (normalization, log2 transformation, quality control) was accepted; downstream processing removed empty probes and multimapping probes and collapsed multiple probes per gene to median expression values. Clinical covariates such as treatment and blood-cell counts were not available in the GEO datasets and therefore were not used as covariates.
The RMVar database of RNA modification–associated variants was cross-referenced with eQTLGen cis-eQTL results and ALS differentially expressed genes (DEGs) from the GEO cohorts. Genes with both RMVar m6A-SNP annotations and cis-eQTL signals were considered m6A-SNP–related genes. Intersecting these with ALS-associated DEGs yielded 109 candidate genes for downstream feature selection.
Gene ontology enrichment (using Metascape) was performed to explore biological processes, molecular functions, and cellular components associated with identified genes; terms with BH-adjusted p < 0.05 were considered significant (detailed results reported in supplementary materials). Differential expression analysis used limma with thresholds |log2FC| > 0.25 and adjusted p < 0.05. Visualization included heatmaps of DEGs.
A combined machine-learning approach screened the 109 candidate genes:
Stability was assessed with 100 stratified bootstrap resampling iterations repeating the full selection pipeline. Selection frequency across iterations informed feature stability. Because a complete nested cross-validation estimate of the entire pipeline was not produced, training-cohort performance is reported as apparent performance; unchanged-coefficient external validation was treated as the primary test of generalizability.
The final diagnostic panel comprised seven genes: TMED5, OXR1, BRI3, FEM1C, SUZ12, EIF2AK4, and TJAP1. Using regression coefficients estimated in the training cohort (GSE112676), a seven-gene model and a nomogram were constructed. Coefficients were applied without re-estimation to the external validation cohort (GSE112680). Receiver operating characteristic (ROC) analysis was used to evaluate discrimination (AUC with 95% CI via DeLong); calibration metrics included calibration intercept and slope, Brier score, and Hosmer–Lemeshow test. Supplementary tables summarize marker-specific and combined-model AUCs and calibration results; recalibrated analyses were provided as sensitivity checks.
Because direct blood-cell counts were unavailable, immune-cell proportions were inferred from transcriptomic data using CIBERSORT with the LM22 signature. The study estimated relative proportions of 22 immune cell types and performed principal component analysis to assess whether inferred immune profiles could distinguish ALS from controls. Significant differences were reported for inferred proportions of monocytes, neutrophils, and T-cell subsets. Correlations between candidate gene expression and inferred immune-cell proportions were evaluated using Spearman correlation.
Sequence-based prediction of m6A modification sites near selected SNP loci was performed using SRAMP. These results are in silico predictions and not experimental evidence of methylation. Selected SNP loci were annotated to nearby RBP-binding regions using available annotation resources, suggesting possible mechanisms by which SNPs could affect RNA–protein interactions and m6A regulation.
An exploratory validation cohort (10 ALS, 10 controls) was analyzed for gene expression and global m6A levels. Findings reported in the source include significant upregulation of FEM1C and SUZ12 at both mRNA and protein levels in ALS peripheral blood and a reduction in global m6A modification levels in ALS samples. Detailed primer sequences and antibodies used for qRT-PCR and Western blot are provided in the source tables.
This multi-omics integration identified a set of m6A-SNP–related candidate markers and a seven-gene diagnostic model that discriminated ALS from controls in the training cohort and retained moderate performance in an external validation cohort. The study links candidate loci to predicted m6A sites and RBP annotations and reports altered inferred immune-cell composition in ALS blood. Limitations noted by the authors include the small exploratory clinical sample size, lack of certain clinical covariates in GEO datasets, absence of a fully nested cross-validation for the selection pipeline, and the need for larger independent cohorts and further calibration before clinical application of the markers.