Perfluorooctanoic acid (PFOA) is a widely used perfluoroalkyl substance (PFAS) with high environmental persistence and bioaccumulation. Epidemiological and experimental evidence has associated PFOA with hepatic injury and disturbances in lipid metabolism, suggesting it may be an environmental risk factor for steatotic liver disease. The 2023 clinical nomenclature shift from NAFLD to Metabolic Dysfunction-Associated Steatotic Liver Disease (MASLD) emphasizes metabolic dysfunction as central to disease definition and motivated a systems-level investigation of environmental contributors such as PFOA.
This study applied an integrated framework—computational toxicology, transcriptomic and single-cell multi-omics, machine learning, molecular docking and dynamics, and in vitro experiments—to interrogate potential molecular and cellular mechanisms linking PFOA exposure to MASLD pathogenesis.
The analytic workflow combined (1) target prediction for PFOA using public chemogenomic resources; (2) identification of MASLD/NAFLD-associated genes from multiple disease databases and transcriptomic datasets; (3) a multi-algorithm machine-learning pipeline to prioritize core genes; (4) single-cell RNA-seq analysis to resolve cell type–specific expression; (5) molecular docking and 100 ns molecular dynamics (MD) to evaluate ligand–protein binding stability; and (6) preliminary in vitro validation in an FFA-induced HepG2 MASLD model.
The 3D structure and canonical SMILES for PFOA were obtained from PubChem (SMILES: C(=O)(C(C(C(C(C(C(C(F)(F)F)(F)F)(F)F)(F)F)(F)F)(F)F)(F)F)O). Target prediction combined ChEMBL, SwissTargetPrediction (probability > 0), and STITCH (interaction score > 0.4), yielding 271 predicted human protein targets. Nineteen NAFLD-related GEO datasets (4 microarray, 1 NGS, 14 scRNA-seq) were downloaded; three microarray datasets (GSE66676, GSE89632, GSE164760) were merged as the training set after probe annotation, normalization, and ComBat batch correction (PCA used to verify correction). GSE63067 and GSE135251 served as independent validation sets.
Differential expression analysis on the merged training set using limma (|log2FC| > 0.585, adjusted P < 0.05) identified 652 DEGs. Disease-related targets were compiled from GeneCards, OMIM, and the Therapeutic Target Database resulting in a nonredundant NAFLD gene library of 902 targets. Intersection of the 271 predicted PFOA targets with the 902 disease genes produced 17 shared targets for downstream prioritization.
To refine candidate genes, the authors implemented a comprehensive machine-learning pipeline involving 11 algorithm types and generated 113 model combinations in a two-stage feature selection plus modeling framework. Algorithms included regularized regressions (Lasso, Ridge, Elastic Net), SVM, LDA, Random Forest, GBM, XGBoost, glmBoost, stepwise logistic regression, PLS, and Naïve Bayes. Hyperparameters were optimized by 10-fold cross-validation (fixed seed 123). Model performance metrics included accuracy, sensitivity, specificity, F1-score, and AUC-ROC.
SHAP (SHapley Additive exPlanations) interpretability analysis was applied to the best-performing model to quantify feature contributions. Through this pipeline, six hub genes were prioritized: NR4A2, BCL6, CASP1, SHBG, FABP4, IL10. Diagnostic performance of models incorporating these genes reached AUC values up to 0.996.
Fourteen scRNA-seq samples (7 healthy, 7 NAFLD) from GEO were analyzed with Seurat v5. Quality control filtered cells with fewer than 200 genes and genes expressed in fewer than 3 cells; additional thresholds were applied for gene counts, mitochondrial percentage, UMI counts, and genes-per-UMI metrics. Doublets were removed using scDblFinder. Harmony integration corrected batch effects across samples; 2,000 highly variable genes were used and scaling regressed out mitochondrial and cell cycle effects.
Clustering employed PCA, a K-nearest neighbor graph, and the Louvain algorithm (resolution 0.8), with UMAP visualization. Cell types were annotated by canonical markers. Proportions of cell types per sample were calculated and compared between healthy and NAFLD groups using Wilcoxon rank-sum tests. Monocle2 pseudotime analysis was applied to macrophage and monocyte subsets to infer trajectory dynamics.
The six hub genes showed distinct expression patterns across liver cell subtypes, indicating cell-type-specific involvement in MASLD-related processes.
Human protein 3D structures were obtained from UniProt and PFOA from PubChem. Molecular docking was performed on the CB-Dock platform; the lowest Vina-score poses for SHBG and FABP4 complexed with PFOA were selected for MD. MD simulations were run in Gromacs2023.2 using the amber14sb force field for proteins, GAFF2 for PFOA, TIP3P water, and standard equilibration (NVT and NPT) followed by 100 ns production runs. Simulations indicated stable binding of PFOA to SHBG and FABP4 over the production trajectory.
HepG2 cells were used to model MASLD by induction with 0.5 mM FFA-BSA (sodium oleate:sodium palmitate 2:1) for 24 hours, with lipid accumulation confirmed by Oil Red O staining. After model establishment, cells were exposed to 80 μg/mL PFOA for 24 hours; control and model groups without PFOA were included. Experiments were performed in biological triplicate (n = 3).
Cell culture reagents, PFOA (purity ≥ 97%), primary antibodies for Western blotting (BCL6, IL10, SHBG, NR4A2, cleaved Caspase-1, FABP4, β-actin), and CCK-8 viability assays were employed. In vitro results showed that PFOA exposure significantly altered mRNA and protein expression levels of the identified core genes in the MASLD HepG2 model.
Collectively, results are consistent with a mechanistic link whereby environmental PFOA exposure may contribute to MASLD pathogenesis through disruption of lipid metabolism, inflammatory responses, and immune regulation.
The authors emphasize that the integrated computational and experimental data identify biologically plausible pathways but do not establish epidemiological causation. Key limitations include reliance on public datasets, predictive target identification methods, and a single in vitro hepatoma cell line model. The study does not provide prospective, quantified human exposure–outcome data. The authors call for prospective studies with measured PFOA exposure to confirm causality in human populations.
In summary, this multi-omics and computational toxicology study highlights six core genes and supports mechanistic plausibility for PFOA involvement in MASLD, offering candidate molecular targets for further investigation and risk assessment.