Hypoglycemia is a common and potentially serious complication of glucose-lowering therapy in patients with diabetes. The authors performed a systematic review and meta-analysis to evaluate the predictive performance, methodological quality, and clinical applicability of machine learning–based models developed to predict hypoglycemia in Chinese patients with diabetes. The review targeted studies conducted in China and searched the literature through February 2026.
A comprehensive search was conducted in PubMed, Embase, Web of Science, the Cochrane Library, CINAHL, CNKI, and Wanfang from database inception to February 2026. The search combined MeSH and free-text terms for hypoglycemia, diabetes, and ML-based prediction models. Only full-text original research articles in English or Chinese reporting AUC as a performance metric and developing ML-based hypoglycemia prediction models in Chinese patients with diabetes were eligible. Observational designs (cohort or case-control) were included. Studies solely reporting risk factors, non-original reports, or with unavailable full text were excluded.
After deduplication and screening, 92 full texts were assessed and 13 studies met the inclusion criteria and were included in the final meta-analysis. Reasons for excluding full-text articles included lack of relevant outcomes, non-Chinese populations, analyses limited to risk factors without model development, or language restrictions. The flow of study selection followed PRISMA-DTA guidance and is summarized by the authors.
Thirteen studies published between 2022 and 2026, all conducted in China, were included. Eleven studies used retrospective designs and two were prospective. Most studies were single-center (11/13) and two used multicenter datasets. The reported prevalence of hypoglycemia across individual studies ranged from 4% to 46.93%; the pooled prevalence (from 12 studies) was 25% (95% CI: 17%–33%). Prediction horizons varied widely from 30 minutes to 12 months, with many models intended for short-term prediction.
A diverse set of algorithms was applied across studies, including logistic regression (LR), decision tree (DT), random forest (RF), support vector machine (SVM), extreme gradient boosting (XGBoost), LightGBM, deep neural networks (DNN), and long short-term memory (LSTM) models. Ensemble methods such as XGBoost and RF were frequently reported as top-performing algorithms in individual studies, with reported AUCs ranging from 0.822 to 0.978.
Sample sizes for model development datasets ranged from 192 to 255,404 participants; validation datasets ranged from 70 to 109,459. Seven studies reported external validation; the remainder used internal validation methods (k-fold cross-validation or bootstrapping). Six studies reported calibration assessment (five used calibration curves and one used the Hosmer–Lemeshow test). No studies reported decision-curve analysis or other explicit measures of clinical utility.
Using a random-effects model, the authors pooled prevalence estimates from 12 studies and estimated an overall hypoglycemia prevalence of 25% (95% CI: 17%–33%). Heterogeneity among prevalence estimates was not detailed here but individual study prevalences varied substantially (4%–46.93%).
The pooled discrimination across the included ML-based models was high, with an overall AUC of 0.90 (95% CI: 0.87–0.93). Subgroup pooled AUCs by algorithm were reported as follows: XGBoost 0.89, RF 0.88, SVM 0.85, LightGBM 0.84, LR 0.83, and DT 0.81. Heterogeneity for many pooled estimates was high (I2 frequently >85%).
Additional subgroup analyses were performed by study design, center type, predictor type, and validation method. Key subgroup results included:
Across the 13 studies, more than 40 unique predictors were reported. Predictors included demographic, clinical, laboratory, treatment, and time-series CGM features. Variables reported in at least two studies and highlighted by the authors included age, insulin use, meal omission, BMI, HbA1c, creatinine, and history of hypoglycemia. The distribution of other predictors varied by study and model type.
Model studies were appraised using PROBAST. Of 13 studies, 4 were judged at low overall risk of bias, 2 unclear, and the remainder at high risk. Most studies had low risk of bias for participant selection and predictors domains, but several studies exhibited issues in outcome definition/assessment and in the analysis domain. Applicability concerns were assessed across participants, predictors, and outcomes; domain-level findings varied between studies. The authors note that methodological limitations, including limited validation and analytic concerns, were common.
The review highlights that research on ML-based hypoglycemia prediction in Chinese patients is at an early stage. Although several models demonstrate good discrimination (pooled AUC 0.90), important limitations were identified: heterogeneous study designs and populations, variable prediction horizons, frequent single-center and retrospective datasets, high between-study heterogeneity, limited external validation, sparse reporting of calibration and no decision-curve or clinical utility analyses, and concerns about robustness and interpretability. The authors recommend further work to develop reliable, externally validated, and interpretable models and to evaluate clinical utility before routine clinical implementation for early risk identification.
Note: This summary is limited to the information reported in the source article. Specific study-level details and individual model metrics are reported in the original tables and supplementary materials cited by the authors.