---
title: "Computational screening of tumor suppressor variants driving gastric cancer pathogenesis"
id: "plos-one-14-computational-screening-of-oncogenic-genetic-variations-in-tumor-suppressor"
canonical_url: "https://medichelpline.com/clinical-feed/plos-one-14-computational-screening-of-oncogenic-genetic-variations-in-tumor-suppressor"
content_type: "clinical_feed_article"
specialty: "Oncology"
source_name: "PLOS ONE (Medicine)"
source_url: "https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0358440"
published_at: "2026-09-17T14:00:00.000Z"
evidence_level: "Journal Feed"
license: "CC-BY-NC-4.0 / Informational Use"
---
# Computational screening of tumor suppressor variants driving gastric cancer pathogenesis
## Provenance & Clinical Metadata
- **Canonical URL:** https://medichelpline.com/clinical-feed/plos-one-14-computational-screening-of-oncogenic-genetic-variations-in-tumor-suppressor
- **Specialty:** [Oncology](https://medichelpline.com/clinical-feed/oncology.md)
- **Primary Source:** PLOS ONE (Medicine)
- **Source URL:** [Original Journal Publication](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0358440)
- **Published At:** 2026-09-17T14:00:00.000Z
- **Evidence Rating:** Journal Feed
## Executive GIST (TL;DR)
- The study used a deep learning-based **graph neural network** to identify genes in gastric cancer (GC) that show both differential expression and high mutation propensity, producing robust model metrics (MSE, R², AUC-ROC reported). - The model prioritized 1,886 genes exhibiting concurrent expression dysregulation and mutation propensity in GC. - Network analysis identified **TP53** as the principal hub gene (47.6% mutation frequency), followed by ERBB2 (8.8%), **CDH1** (8.2%), and **APC** (6.8%). - Computational analyses focused on tumor suppressors **TP53**, **CDH1**, and **APC**, assessing evolutionary conservation, biophysical energetics, and conformational change using unsupervised machine learning. - Validation using cBioPortal confirmed 36 missense SNPs affecting post-translational modification sites (methylation and phosphorylation) and 60 nonsense SNPs associated with GC. - TP53, CDH1, and APC were significantly upregulated in GC tissues and their expression associated with altered patient survival (p < 0.05). - The transcription factor EZH2 and miRNA miR-129-5p emerged as shared regulatory elements influencing all three tumor suppressors. - Mutations in these tumor suppressors correlate with dysregulation of oncogenes including CCNE1/2 and FGFR2, highlighting pathway-level effects. - The study provides a systems-level molecular framework linking pathogenic variants in tumor suppressors to compromised tumor-suppressive function and identifies candidate biomarkers and targets for precision intervention in GC. - All data and supporting information are reported within the paper and its supporting files; no external funding or competing interests were declared.
## Clinical Analysis & Structured Key Points
Computational screening of oncogenic genetic variations in tumor suppressor proteins driving gastric cancer pathogenesis | PLOS One Browse Subject Areas ? Click through the PLOS taxonomy to find articles in your field. For more information about PLOS Subject Areas, click here . Article Authors Metrics Comments Media Coverage Peer Review Reader Comments Figures Figures Abstract Gastric cancer (GC) is currently the fifth most common cancer globally, often driven by dysregulation of tumor suppressor pathways. While individual studies on genetic variations of proteins are common, a comprehensive systems-level analysis of proteins regulating GC pathways showing both expression dysregulation and high mutation frequency remains unexplored. Therefore, our study aimed to identify critical genes and their pathogenic variations disrupting the tumor-suppressive capacity of the GC pathway. We employed a deep learning-based graph neural network model to identify genes exhibiting both dysregulated expression and high mutation propensity in GC. Key tumor suppressor proteins (TP53, CDH1, and APC), their genetic variants, and molecular components were subjected to in-depth computational analyses, including evolutionary conservation profiling, biophysical energetics assessment, and unsupervised machine learning for conformational change detection. Variants were validated from the cBioPortal database and patient survival data. Our deep learning model demonstrated exceptional performance (MSE: 0.00482 ± 0.00023 to 0.07108 ± 0.00437; R²: 0.85098 ± 0.01903 to 0.85899 ± 0.01987; AUC-ROC: 0.93095 ± 0.01758 to 0.93309 ± 0.00725) and identified 1,886 genes exhibiting both differential expression and mutation propensity. Graph neural network analysis revealed TP53 as the most prominent hub gene (47.6% mutation frequency), followed by ERBB2 (8.8%), CDH1 (8.2%), and APC (6.8%). Variants validation from cBioPortal confirmed the association of GC with 36 missense SNPs that critically affect post-translational modification (methylation and phosphorylation) sites and 60 nonsense SNPs. Furthermore, TP53, CDH1, and APC were significantly upregulated in GC tissues and associated with altered patient survival (p < 0.05). The transcription factor EZH2 and miRNA miR-129-5p were identified as key regulatory elements affecting all three tumor suppressors. Additionally, mutations trigger dysregulation of multiple common oncogenes, including CCNE1/2 and FGFR2. This systems-level analysis provides a molecular framework demonstrating how pathogenic variants fundamentally compromise tumor suppressor proteins in GC pathways, leading to the identification of potential biomarkers and precise therapeutic decisions for GC intervention. Citation: Mim FF, Sumiya TA, Khanam R, Akter J, Arzu F, Haque S, et al. (2026) Computational screening of oncogenic genetic variations in tumor suppressor proteins driving gastric cancer pathogenesis. PLoS One 21(9): e0358440. https://doi.org/10.1371/journal.pone.0358440 Editor: Mohammed S. Razzaque, The University of Texas Rio Grande Valley, UNITED STATES OF AMERICA Received: September 26, 2025; Accepted: September 1, 2026; Published: September 17, 2026 Copyright: © 2026 Mim et al. This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Data Availability: All relevant data are within the paper and its Supporting Information files. Funding: The author(s) received no specific funding for this work. Competing interests: The authors have declared that no competing interests exist. 1. Introduction Gastric cancer (GC) is recognized as the primary epithelial malignancy of the stomach, characterized by its complexity and diversity [ 1 – 3 ]. Although there has been a general decline in both incidence and mortality rates in many countries over the past few decades, GC remains the fifth most common cancer and the fourth leading cause of cancer-related deaths globally [ 4 , 5 ]. GC involves dysregulation of several key signaling pathways that contribute to its development, progression, metastasis, and therapeutic resistance, including receptor tyrosine kinases (RTKs) such as EGFR, HER2, VEGFR, and c-MET, which activate downstream pathways like MAPK/ERK and PI3K/AKT/mTOR, promoting cell growth, proliferation, survival, and angiogenesis [ 6 , 7 ]. The p53 pathway, crucial for DNA repair, cell cycle control, and apoptosis, is often disrupted in GC, while the Wnt/β-catenin pathway plays a role in cell proliferation and migration. Additionally, the NF-κB and TGF-β pathways are involved in inflammation and immune response [ 8 ]. In general, tumor suppressor proteins act as guardians of the genome, regulating vital cellular processes like cell cycle control, apoptosis, and DNA damage repair [ 9 ]. Genetic variations in tumor suppressor genes, such as TP53 , CDH1 , and APC, are closely associated with GC pathways and extensively reported in GC patients [ 10 – 14 ]. Tumor suppressor proteins, which normally regulate cell growth and prevent tumor formation, are mutation-prone in most of the cancers, including GC. Pathogenic genetic variations in these proteins can disrupt their normal function, leading to uncontrolled cell growth and tumor development [ 15 , 16 ]. Among genetic variations, frameshift, missense, nonsense, splice site, ncRNA, and UTR (untranslated region) variations are most frequent and predominantly found in humans. Generally, frameshift mutations disrupt the reading frame, missense mutations alter amino acids, nonsense mutations introduce premature stop codons, splice site mutations affect RNA splicing, mutations in ncRNAs impact gene regulation, and UTR mutations influence mRNA stability and translation, all potentially leading to altered or non-functional proteins and various phenotypic consequences, including cancers [ 17 , 18 ]. However, the impact of missense SNPs (Single-nucleotide polymorphisms) on protein function, stability, and disease development is not always straightforward to predict, like other genetic variations. This complexity arises from the intricate interplay between the altered amino acid, its location within the protein structure, its role in the protein’s overall function, and molecular interference by TFs (transcription factors), miRNAs, methylation, phosphorylation, ubiquitylation sites, etc. [ 19 ]. Missense SNPs in tumor suppressor proteins can alter protein structure, stability, and function, thereby disrupting their ability to suppress tumorigenesis [ 20 ]. In addition, missense SNPs can affect tumor suppressor proteins through multiple mechanisms, including altering protein folding and stability, disrupting critical functional domains, and regulating transcriptional, post-transcriptional, and post-translational modifications [ 21 , 22 ]. Evaluating the impact of clinically significant missense SNPs in tumor suppressor proteins and understanding the functional consequences and molecular interferences of missense SNPs is essential for unraveling the complex genetic landscape of GC. Furthermore, these structural and functional changes as a result of the pathogenic genetic variations can impair protein-protein interactions, disrupt signaling pathways, and ultimately impair the tumor suppressor activity, rendering cells more susceptible to malignant transformation and ultimately contributing to the development and progression of cancer. Therefore, elucidating the impact and molecular mechanism of clinically significant pathogenic variations in tumor suppressor proteins TP53, APC, and CDH1 of the GC pathway is crucial for improving our understanding of the genetic factors underlying GC risk and pathogenesis. This knowledge can guide the development of personalized diagnostic and therapeutic strategies, ultimately enhancing patient outcomes. While previous studies have explored the molecular basis of early diagnostic and prognostic biomarkers for gastric cancer [ 23 ] and the effects of missense SNPs on individual tumor suppressor proteins TP53, APC, and CDH1 [ 24 – 26 ], comprehensive analysis to identify key proteins that are highly dysregulated and have high mutation frequency in GC at the systems level is lacking. Additionally, most of the previous studies focus on the identification of an isolated type of pathogenic variation, whereas our study focuses on the identification and validation of the direct impact of all clinically significant pathogenic genetic variations on key proteins that disrupt the balance of the tumor suppressor capability of the GC pathway. Our study also focused on identifying genetic components crucial to disrupting this balance, contributing to aiding this disruption, and corroborating the development of GC, unveiling the molecular mechanisms of genetic variations in GC pathogenesis. The overview of this study is presented in Fig 1 . Download: PNG larger image TIFF original image Fig 1. The sequential flow diagram of the study to establish a link between differentially expressed high-frequency mutation-prone genes, pathogenic genetic variations, and GC pathogenesis. https://doi.org/10.1371/journal.pone.0358440.g001 2. Methodology 2.1. Screening of differentially expressed mutation-prone genes in GC We developed a deep learning prioritization and ranking Graph Neural Networks (GNNs) model, specifically GCNConv (Graph Convolutional Networks), to identify GC pathway-specific differentially expressed mutation-prone genes. At first, RNAseq gene expression data were retrieved from the GEPIA2 database ( http://gepia2.cancer-pku.cn/ ) [ 27 ], and GC-specific mutation-prone genes were retrieved from the cBioPortal database ( https://www.cbioportal.org/ ) [ 28 ]. The common genes between dysregulated genes and mutation-prone genes were identified using a Venn diagram from the InteractiVenn tool ( https://www.interactivenn.net/ ) [ 29 ]. The protein-protein interaction network was constructed using the common genes from the STRING database ( https://string-db.org/ ) [ 30 ], which contains both experimentally validated and predicted protein-protein interactions. Furthermore, GC pathway-specific genes were collected from the KEGG pathway database ( https://www.kegg.jp/ ) [ 31 ] that contains pathways of molecular mechanisms for disease development, including cancer pathways [ 32 ]. These commonly mutation-prone differentially expressed genes (DEGs), genes from protein-protein interaction networks, and GC pathway-specific genes were utilized as input to develop a robust computational framework to identify pathway-specific, highly dysregulated, mutation-prone genes responsible for GC pathogenesis. Secondly, the NetworkX library of Python was implemented for graph construction and topological analysis (degree, betweenness, closeness) [ 33 ]. Thirdly, the PyTorch library for tensor operations and neural network architecture, and torch_geometric for graph neural network implementation, were utilized. Finally, the StandardScaler of the scikit-learn library was used for data preprocessing [ 34 ]. Specifically, we employed a Graph Convolutional Network (GCN) implemented through the GCNConv layers, which enabled the propagation of node features across the network topology [ 35 , 36 ]. Node features were normalized using StandardScaler to ensure consistent scaling across heterogeneous molecular data types [ 34 ]. The network architecture consisted of two GCNConv layers followed by ReLU (Rectified Linear Unit) activation functions to introduce non-linearity into the neural network, which is crucial because most of the real-world problems are non-linear, and neural networks need non-linear activation functions to learn complex patterns [ 37 ]. Subsequently, dropout regularization (rate = 0.3) was applied between layers to prevent overfitting. Model performance was rigorously assessed through train-validation-test splitting (60:20:20 ratio) using the train_test_split module from scikit-learn, where training and validation datasets were utilized for learning and internal validation purposes, whereas the testing set was utilized and masked for external validation. Moreover, a nested 5-fold cross-validation (CV) strategy was applied to maintain both hyperparameter tuning and unbiased test estimate with 2 loops: outer loop for model evaluation and inner loop for hyperparameter tuning. In this framework, the inner loop was exclusively used to optimize model hyperparameters without exposing the outer test folds, whereas the outer loop evaluated performance on completely unseen data partitions to provide an unbiased estimate of generalization performance. To mitigate information leakage, we implemented a gene-level partitioning strategy in which all features and network relationships associated with a given gene are confined to a single data split, preventing cross-set propagation of topological or annotation-derived information. Finally, the performance metrics of the model were evaluated using training and validation loss over 200 epochs. To develop the model, the GNN is first trained to predict a propensity score based on the relevance of 3 gene categories. (1) Mutation-prone Differentially Expressed Genes (FDR-adjusted p-values < 0.05). (2) Genes from the protein-protein interaction network (confidence score: 0.4 (medium to balanced network). (3) Gastric cancer (GC) pathway-specific genes. This represents the integrated likelihood that a gene contributes to gastric cancer pathogenesis due to mutations, expression dysregulation, and network-level functional importance. A continuous propensity score is operationalized by combining normalized gene expression magnitude (Log2FC) with multiple graph-topological centrality measures (degree, betweenness, and closeness) and, when available, pathway membership information. These biological components reflect the premise that genes that are both dysregulated and centrally positioned in protein interaction networks, along with being present in the GC pathway, are more likely to act as key disease drivers. The resulting composite score serves as the dependent variable in a regression framework and ranks the top genes. Therefore, the primary regression performance metrics were mean squared error (MSE), root mean squared error (RMSE), and coefficient of determination (R²). Afterward, the classification metrics were used as a secondary interpretive analysis (AUC-ROC reporting) rather than the primary modeling objective. For this purpose, the continuous propensity scores were binarized using a biologically interpretable threshold (≥0.5 after sigmoid normalization), representing high-confidence versus low-confidence pathogenic candidates. This transformation allowed us to assess the model’s ability to discriminate high-risk oncogenic genes from background genes, resulting in identifying the top 20 hub genes based on both regression and classification metrics. Due to the complexity of several layers of input data in the model, we did not consider the Youden index or cost-sensitive approaches in this case; rather, we chose a simplified classification boundary threshold (≥0.5 after sigmoid normalization) to remove background genes from more valuable genes. The regression captured fine-grained biological risk ranking, while ROC and precision-recall analyses provided the discrimination performance to remove background genes for prioritization tasks. All reported metrics were computed on independent training, validation, and test partitions to ensure robustness. 2.2. Identification of high-frequency mutation-prone genes We have screened high-frequency mutation-prone genes by analyzing 147 GC patient samples from the cBioPortal database due to their impact on the dysregulation of proteins. Based on the deep learning neural network model and frequency analysis, we have found that tumor suppressor proteins are the most frequently mutated proteins. We have further examined their involvement in the GC pathway from the KEGG database ( https://www.genome.jp/kegg/ ), which is a comprehensive resource for understanding high-level biological functions and utilities, including molecular-level information such as genes, proteins, and biochemical pathways [ 31 ]. Another crucial reason for choosing tumor suppressor proteins for single-nucleotide polymorphism analysis is that they play critical roles in inhibiting GC, and mutations in these genes are likely to have been implicated in the development and progression of GC [ 15 , 16 ]. 2.3. Identification of tumor suppressor domains in proteins The tumor suppressor domains of proteins were identified using the NCBI Conserved Domain Search Tool ( https://www.ncbi.nlm.nih.gov/Structure/cdd/wrpsb.cgi ) [ 38 ] and validated through the UniProt database ( https://www.uniprot.org/ ) [ 39 ]. The NCBI Conserved Domain Search Tool was used to search for conserved domains within the target protein sequences. This allows searching against various domain databases, including the NCBI-curated Conserved Domain Database (CDD), to identify significant matches to known functional domains. The UniProt database was then used to gather additional information on the identified tumor suppressor domains responsible for tumor suppressor activity. 2.4. Retrieval of clinical variations of tumor suppressor proteins Clinically significant variations were retrieved from the ClinVar database of NCBI ( https://www.ncbi.nlm.nih.gov/clinvar/ ) to investigate their association with cancer [ 40 ]. ClinVar consists of curated and annotated clinically significant variants, such as frameshift, missense, nonsense, splice site, ncRNA, and UTR, to prioritize potentially pathogenic variations and provide evidence of a link between these variants and disease phenotypes. Only pathogenic variations were selected, and likely pathogenic clinical variations were considered only for missense SNP variation analysis, as sophisticated methodologies have been established to analyze missense variations compared to other variations, and these variants are often directly linked to functional consequences [ 41 ]. 2.5. Identifying damaging, deleterious, and disease-causing missense SNPs To identify potentially damaging, deleterious, disease-causing, and oncogenic missense SNPs, we employed a multi-faceted bioinformatics approach utilizing seven established prediction tools. We leverage the capabilities of PolyPhen2 ( https://genetics.bwh.harvard.edu/pph2/ ) [ 42 ], SIFT ( https://sift.bii.a-star.edu.sg/ ) [ 43 ], SNPs&GO ( https://snps.biofold.org/snps-and-go/ ) [ 44 ], PhD-SNP ( https://snps.biofold.org/phd-snp/ ) [ 45 ], SNPs3D ( http://snps3d.org/ ) [ 46 ], GVGD ( http://agvgd.hci.utah.edu/ ) [ 47 ], and Cscape ( http://cscape.biocompute.org.uk/ ) [ 48 ]. Each tool was selected for its unique strengths to meet the ACMG variant classification guidelines [ 49 ], which are standard practice in SNP studies [ 50 , 51 ]. PolyPhen2 assessed the potential impact of amino acid substitutions on protein function using structural and comparative considerations, whereas SIFT predicted the effects of amino acid changes based on sequence homology and physical properties. SNPs&GO correlated SNPs with gene ontology to predict disease relevance, PhD-SNP evaluated the deleterious effects of mutations, SNPs3D provided structural insights into damaging SNPs, GVGD classified the likelihood of pathogenicity, and finally, Cscape evaluated the oncogenic potential of mutations. Incorporating the scores of seven computational tools improved forecast precision and ensured the stringency and accuracy of results. 2.6. Determining the impact of missense SNPs on protein structural stability To determine the impact of missense SNPs on the structural stability of tumor suppressor proteins, we utilized the I-Mutant 2.0 tool [ https://folding.biofo
## Related Clinical Research

- [Lived Experience in Rare Disease Research: Shaping Outcomes for Patients](https://medichelpline.com/clinical-feed/plos-medicine-0-living-with-a-rare-disease-why-lived-experience-must-shape-research.md)
- [Early 68Ga-PSMA-11 PET/CT to Predict 6-Month Disease Control in Metastatic ccRCC: PSMA-RENAL Phase](https://medichelpline.com/clinical-feed/bmj-open-18-can-early-68ga-psma-11-pet-ct-predict-6-month-disease-control-in-patients.md)
- [CSN1 and the COP9 signalosome regulate glioblastoma stem cell growth and tumorigenesis](https://medichelpline.com/clinical-feed/plos-one-4-the-cop9-signalosome-subunit-csn1-regulates-glioblastoma-stem-cell-growth.md)
- [Blood test trends versus single abnormal results to detect cancer in patients with unexpected weig](https://medichelpline.com/clinical-feed/plos-medicine-1-blood-test-trend-versus-single-threshold-abnormality-to-discriminate-cancer.md)
- [Financial toxicity in families of children with cancer: qualitative systematic review and meta-agg](https://medichelpline.com/clinical-feed/pubmed-42763337.md) (DOI: 10.1007/s00520-026-11232-6)

## Navigation
- [← Back to Oncology Feed](https://medichelpline.com/clinical-feed/oncology.md)
- [← All Clinical Specialties](https://medichelpline.com/clinical-feed.md)
## Medical & Regulatory Disclaimer

> [!CAUTION]
> MedicHelpline content is structured for research, educational, and professional discovery purposes. It does not constitute individual medical advice, clinical diagnosis, or treatment recommendations.
> Always verify dosing, contraindications, and regulatory alerts against official product labeling and primary regulatory sources before clinical decision-making.