This study used nationally representative survey data to identify multidimensional cardiovascular risk profiles among adults with self-reported rheumatoid arthritis (RA). The authors sought to determine whether combined patterns of metabolic, inflammatory, renal, behavioral, and psychosocial factors form distinct phenotypes and how those phenotypes associate with prevalent cardiovascular disease (CVD).
The analysis included 3,252 adults aged over 20 years with self-reported RA drawn from the 2005–2018 National Health and Nutrition Examination Survey (NHANES). The dataset is publicly available from the National Center for Health Statistics (NCHS), Centers for Disease Control and Prevention (CDC). The source reports that all data used are accessible without restriction. Details on specific variable coding and exact inclusion/exclusion steps were not reported in the source beyond these high-level sample criteria.
Unsupervised partitioning around medoids (PAM) clustering was performed using Gower distance to accommodate mixed variable types. Cluster validity was evaluated using silhouette analysis and internal train–test validation to assess separation and internal correspondence. Associations between cluster (phenotype) membership and prevalent CVD were estimated with survey-weighted binary logistic regression, adjusting for relevant covariates as described in the source. Model discrimination for prevalent CVD was reported as area under the receiver operating characteristic curve (AUC = 0.72).
Six clinically interpretable phenotypes emerged from the clustering analysis. The source reported the relative prevalence estimates for these clusters and labeled them as follows:
The labels reflect dominant patterns of cardiometabolic risk factors, demographic characteristics, and behavioral features captured by the clustering algorithm. The source emphasizes that the phenotypes represent combined patterns across metabolic, inflammatory, renal, behavioral, and psychosocial domains, though the source did not enumerate every variable that drove each cluster in the summary provided.
Using the Cardiometabolic phenotype as the reference group, the study reported the following adjusted associations with prevalent CVD:
These results indicate the highest odds of prevalent CVD among participants assigned to the Severe Metabolic Diabetic and Aging Diabetic Hypertensive phenotypes. The source reported these associations using survey-weighted logistic regression; the full set of covariates included in adjusted models was not detailed in the summary provided.
The six-cluster solution demonstrated moderate separation by silhouette metrics and strong internal correspondence across training and testing sets according to the authors. A phenotype-informed model that included cluster membership showed moderate discrimination for prevalent CVD with an AUC of 0.72. These performance indicators suggest the phenotypes capture meaningful heterogeneity in cardiovascular risk among adults with RA within the NHANES sample.
The findings support the concept that cardiovascular risk in RA is heterogeneous and multidimensional. Identifying data-driven phenotypes that combine metabolic, inflammatory, renal, behavioral, and psychosocial features may provide a more comprehensive framework for characterizing risk than single-factor approaches. According to the source, phenotype-based stratification could inform targeted risk assessment and potentially guide prevention strategies, though translation into clinical pathways would require further validation and operationalization beyond the survey analysis.
The authors state the underlying NHANES data are publicly available from the NCHS/CDC repository. The preprint declares no competing interests and confirms adherence to relevant ethical guidelines; IRB/oversight information and participant consent statements are reported as provided in the source. The source did not provide additional licensing or implementation guidance beyond these statements.
Note: The rewritten summary reports only results and details as presented in the source material. Where the source did not provide specific methodological granularities (for example, the exact variable set driving each cluster or the full covariate list used in adjusted models), those details were not added or inferred.