This study evaluated whether commonly recorded clinical and demographic information can predict fluoroquinolone (FQ) resistance among patients with rifampicin-resistant or multidrug-resistant tuberculosis (RR/MDR-TB). Investigators used data from the TB Portals platform comprising 5,175 patients with RR-TB and available FQ drug susceptibility testing (DST) from eight countries (Azerbaijan, Belarus, Georgia, Kazakhstan, Kyrgyzstan, Moldova, Romania, Ukraine) collected between 2012 and 2024. In the pooled dataset, 1,772 patients (34.2%) had FQ-resistant TB.
Models were developed using three algorithms—logistic regression, neural networks, and XGBoost—and evaluated under a pooled multi-country training strategy. After correcting for optimism, pooled models achieved moderate discrimination: area under the receiver operating characteristic curve (AUROC) values between approximately 0.70 and 0.72, and area under the precision-recall curve (AUPRC) values around 0.57 to 0.59 across the three algorithms. These results indicate modest ability to distinguish FQ-resistant from FQ-susceptible RR-TB cases when training data combine multiple countries.
The authors also fitted models separately within individual countries and assessed performance using internal validation. Within-country models tended to perform better than pooled models in several settings. In some countries, AUROC and AUPRC values reached about 0.8, demonstrating that locally trained models can achieve substantially higher discrimination than models trained on pooled international data. This finding suggests that country-specific patterns in predictors and resistance prevalence can improve model fit and predictive performance when models are developed and validated within the same setting.
To test model generalizability, the study implemented cross-country external validation: models were trained on subsets of countries and evaluated on held-out countries not used in model development. External validation showed heterogeneous results. For certain held-out countries, the loss in performance relative to within-country or pooled validation was negligible, indicating reasonable transferability in some contexts. For other countries, however, the decrease in AUROC or AUPRC exceeded 0.1, signifying substantial performance degradation when applying externally trained models.
These variable outcomes underscore that a model developed in one country or group of countries cannot be assumed to perform similarly in another without rigorous external validation. The magnitude of transferability loss depended on both the target country and the modeling algorithm.
Across modeling approaches and validation strategies, a limited set of predictors were consistently informative. Variables related to case definition and prior treatment history were among the most reliable predictors of FQ resistance. By contrast, demographic factors, comorbidities, social-risk indicators, education, and employment variables showed more variable importance across countries and algorithms. The inconsistent contribution of these latter predictors likely contributed to differences in model performance between countries and to the variable external validity observed in cross-country testing.
Data were drawn from TB Portals, an open-access TB data resource curated by the National Institute of Allergy and Infectious Diseases. The analytic sample included 5,175 RR-TB patients with FQ DST results from eight countries in Eastern Europe and Central Asia, with collection years spanning 2012–2024. The prevalence of FQ resistance in this sample was 34.2% (1,772 of 5,175 patients).
The authors report that the incidence of RR-TB and FQ resistance in the dataset was relatively stable over the study period; they note this stability as a context for interpreting model performance and generalizability.
The authors emphasize several limitations derived from the source dataset and study design. One key limitation is the relative stability of RR-TB and FQ resistance incidence in the analyzed data; because the dataset did not reflect marked temporal shifts in MDR-TB dynamics, the findings may not generalize to settings experiencing rapid changes in resistance patterns or to periods with substantially different prevalence of resistance.
Additional limitations implicit in the study design include reliance on predictors available in the TB Portals dataset and on retrospective DST results. The study did not include rapid molecular DST availability as a variable and did not attempt to replace laboratory testing. The study was not registered and no external protocol was prepared, and the authors explicitly acknowledge these procedural points.
Prediction models using readily available clinical and demographic information can provide some indication of FQ resistance among patients with RR-TB when laboratory DST is unavailable. However, across algorithms and validation strategies, model accuracy was generally moderate, and performance varied substantially by country.
Consequently, the authors conclude that prediction tools developed in one country or a set of countries should not be assumed to generalize to other settings without careful external validation. The most consistent predictive signals were related to case definition and treatment history, suggesting these variables are helpful when constructing locally adapted risk models.
The findings support two practical recommendations drawn from the study data: first, pursue development and validation of locally informed prediction models if clinical prediction is to be used to guide empiric regimen selection; second, prioritize expanded access to rapid and comprehensive drug susceptibility testing so that antibiotic selection for RR/MDR-TB is guided by laboratory evidence rather than prediction alone.