The authors reanalyzed a publicly available hyperlipidemic liver bulk-transcriptomic dataset (GSE338111) to examine how inter-individual transcriptional heterogeneity affects detection of treatment-associated molecular responses. Rather than relying solely on treatment labels, they applied a PC1-guided transcriptomic stratification approach to capture major axes of variation across samples before contrasting groups.
A standard, sex-adjusted differential expression analysis comparing Amlexanox-treated versus DMSO-treated samples in GSE338111 identified only 21 differentially expressed genes at FDR < 0.05 with |log2 fold change| ≥ 1. Submission of this small DEG list to Metascape produced no enriched GO Biological Process terms, indicating limited functional resolution from the conventional analysis in this cohort. The authors interpreted these results as potentially reflecting that animals under the same experimental condition can show divergent molecular responses that mask treatment effects when analyzed solely by treatment group.
To address cohort heterogeneity, the investigators selected the 500 most variable genes across samples and performed principal component analysis. They used the first principal component (PC1) to guide an unsupervised stratification of samples into groups independent of treatment assignment. This strategy prioritizes dominant patterns of transcriptional variability that may reflect biological substructure, cellular composition differences, or response heterogeneity.
PC1-guided clustering resolved three distinct groups within the cohort, denoted G1, G2, and G3. Partitioning samples by PC1 rather than treatment labels revealed transcriptional subgroups that were not captured by the original treatment-based comparison. The resolved groups formed the basis for targeted contrasts, enabling detection of signal associated with subgroup-specific transcriptional states.
From the contrast between G2 and G1, the authors derived a myeloid-associated 20-gene signature. This 20-gene set represents genes whose expression differentiates the PC1-defined groups and, based on downstream analyses, appears linked to myeloid-related transcriptional programs in the liver. The signature was proposed as a compact molecular readout of the PC1-resolved myeloid-associated state.
The study evaluated the signature in independent bulk-transcriptomic cohorts and found that signature expression responded to both dietary challenge and pharmacologic intervention in those datasets. These cross-cohort responses support the signature's reproducibility and sensitivity to metabolic perturbations and treatment, suggesting potential utility for comparing interventions across studies despite cohort heterogeneity.
Single-cell transcriptomic analysis was used to localize the cellular source of the signature. Expression of the 20-gene signature mapped predominantly to hepatic myeloid populations, providing cellular-level evidence that the bulk-tissue signature reflects myeloid compartment transcriptional activity rather than hepatocyte-specific programs.
To probe translational relevance, the authors conducted human cis-eQTL Mendelian randomization and colocalization analyses for genes in the signature. Among signature genes, TAGLN2 emerged with the strongest genetic support for an association with coronary heart disease based on the reported cis-eQTL MR and colocalization results. The source reports TAGLN2 as the gene with the most robust genetic evidence in these analyses.
The authors conclude that PC1-guided stratification improves resolution of heterogeneous hepatic transcriptional responses in cohorts where conventional treatment-based analyses yield limited results. By resolving PC1-defined subgroups and deriving a reproducible myeloid-associated 20-gene signature, this approach can generate cross-cohort molecular markers suitable for further mechanistic studies and translational evaluation. The single-cell localization and human genetic support for TAGLN2 highlight potential directions for follow-up work.
Note: The source article reported these methods, groupings, gene counts, cohort validation and genetic analyses; detailed gene lists, effect sizes, statistical parameters beyond those stated, and specific cohort identifiers for cross-cohort validation were not provided in the source summary and therefore are not included here.