High-risk age at childbirth—defined in this study as childbirth at ≤18 years or ≥40 years—remains an important public health concern in Bangladesh. Early and advanced maternal ages are associated with adverse perinatal and maternal outcomes, including low birth weight, preterm delivery, and increased obstetric complications. In Bangladesh, persistent patterns of early marriage and rural–urban disparities contribute to early childbearing, with prior BDHS rounds reporting low median ages at first marriage and first birth.
This study aimed to identify the socio-demographic and behavioral determinants of high-risk age at childbirth and to evaluate the predictive performance of multiple machine learning algorithms using the nationally representative BDHS 2022 dataset. The authors sought to provide an evidence-based analytical framework to inform targeted interventions for improving maternal and child health outcomes.
Data were drawn from the Bangladesh Demographic and Health Survey (BDHS 2022), a nationally representative survey that uses a two-stage stratified random sampling approach to cover all administrative divisions. The BDHS collects repeated measures on demographic and health indicators, allowing assessment of factors such as maternal and partner education, contraceptive use, and media exposure. The analysis focused on ever-married women aged 15–49; never-married women were excluded. No new data were collected for this study.
The initial dataset comprised 30,078 observations of ever-married women. After cleaning, removing duplicates, and excluding records with missing values, the analytic sample consisted of 15,386 women. The authors used a complete-case approach and did not apply imputation methods, citing concerns that imputation could introduce additional assumptions.
Three feature-selection methods were applied to identify candidate predictors for model training: Lasso (yielding 12 features), Chi Square (13 features), and Boruta (10 features). Selected socio-demographic and behavioral variables considered included maternal age, early marriage indicators, contraceptive use, husband occupation, number of children, husband age and education, respondent age, regional variables, and media exposure.
The cleaned dataset was split into training (80%) and testing (20%) subsets. To address class imbalance, the Synthetic Minority Over-sampling Technique (SMOTE) was applied during training. Models were trained using 10-fold stratified cross-validation. A range of machine learning classifiers were evaluated:
Model performance was assessed using accuracy, precision, recall, F1 score, Area Under the Receiver Operating Characteristic Curve (AUROC), and Area Under the Precision-Recall Curve (AUPRC). Reported results by feature-selection method included:
Lasso-selected features (12 features): CART showed the best performance with Accuracy = 0.77, F1-score = 0.69, AUROC = 0.74, AUPRC = 0.54.
Chi Square-selected features (13 features): GBM achieved the highest performance with Accuracy = 0.80, F1-score = 0.71, AUROC = 0.79, AUPRC = 0.64.
Boruta-selected features (10 features): GBM again performed strongly with Accuracy = 0.75, F1-score = 0.64, AUROC = 0.69.
Overall, the GBM model with Chi Square-selected features demonstrated the strongest predictive performance across reported metrics, achieving the highest AUROC (0.79) and F1-score (0.71).
Across models and feature-selection strategies, the most influential predictors included maternal age, indicators of early marriage, current contraceptive use, husband occupation, number of children, husband age and education, respondent age, and regional disparities. The analysis also highlighted the additional relevance of partner education and media exposure beyond established maternal-level factors.
The study’s findings align with prior research from South Asia and other low- and middle-income country contexts that link educational attainment, marital timing, socioeconomic status, and access to information with timing of first birth. Machine learning approaches in this analysis allowed systematic comparison of algorithms and feature sets and identified tree-based ensemble methods—particularly GBM—as top performers for predicting high-risk ages at childbirth.
From a policy perspective, the authors suggest interventions that target determinants identified by the models: delaying age at first birth, expanding female education, improving contraceptive access and use, and focusing resources on high-risk districts and population subgroups. Emphasizing partner education and mass-media strategies may also support behavioral and normative shifts that reduce early childbearing.
The study used a complete-case analytic approach and did not perform data imputation; as a result, findings reflect only respondents with complete records. The cross-sectional nature of BDHS data limits causal inference. Details about some methodological choices (for example, the full list of variables measured and preprocessing steps) are reported in the source but any additional context beyond the described procedures was not presented here.
Using BDHS 2022 data and multiple machine learning algorithms, this study identified key socio-demographic and behavioral determinants of high-risk age at childbirth in Bangladesh and demonstrated that ensemble tree-based methods, particularly GBM with Chi Square-selected features, provided the best predictive performance. The results offer an evidence-based framework to guide policymakers in designing targeted interventions—such as delaying first birth and expanding female education—to improve maternal and child health outcomes.