---
title: "Machine learning and deep learning prediction of hypertension and key risk factors in Bangladesh"
id: "plos-one-12-machine-learning-and-deep-learning-based-prediction-of-hypertension-and"
canonical_url: "https://medichelpline.com/clinical-feed/plos-one-12-machine-learning-and-deep-learning-based-prediction-of-hypertension-and"
content_type: "clinical_feed_article"
specialty: "Cardiology"
source_name: "PLOS ONE (Medicine)"
source_url: "https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0358471"
published_at: "2026-09-17T14:00:00.000Z"
evidence_level: "Journal Feed"
license: "CC-BY-NC-4.0 / Informational Use"
---
# Machine learning and deep learning prediction of hypertension and key risk factors in Bangladesh
## Provenance & Clinical Metadata
- **Canonical URL:** https://medichelpline.com/clinical-feed/plos-one-12-machine-learning-and-deep-learning-based-prediction-of-hypertension-and
- **Specialty:** [Cardiology](https://medichelpline.com/clinical-feed/cardiology.md)
- **Primary Source:** PLOS ONE (Medicine)
- **Source URL:** [Original Journal Publication](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0358471)
- **Published At:** 2026-09-17T14:00:00.000Z
- **Evidence Rating:** Journal Feed
## Executive GIST (TL;DR)
- This study used the 2022 Bangladesh Demographic and Health Survey (BDHS) cross-sectional data on 14,283 adults (≥18 years) to estimate **hypertension** prevalence and to develop predictive models using machine learning (ML) and deep learning (DL) approaches. - Overall hypertension prevalence was 18.04% (95% CI: 17.2%–18.9%), higher in women (18.87%) than men (16.97%). - Chi-square testing indicated significant associations between hypertension and **age**, **body mass index (BMI)**, **diabetes**, **wealth index**, **education**, **household size**, and **region** (p < 0.05). - Six predictive models were evaluated: four ML models (weighted logistic regression [WLR], random forest [RF], extreme gradient boosting [XGBoost], light gradient boosting machine [LightGBM]) and two DL models (TabNet, multi-layer perceptron [MLP]). - Performance metrics included accuracy, precision, recall (sensitivity), specificity, F1-score, AUC-ROC, and AUC-PR. - Weighted logistic regression achieved the highest accuracy (0.817), precision (0.444), specificity (0.981), AUC-ROC (0.751), and AUC-PR (0.357) on test data but had very low recall (0.070), limiting its sensitivity for detecting hypertensives. - Random forest produced the highest recall (0.687) and highest F1-score (0.460), indicating better sensitivity and more balanced performance for identifying people with hypertension. - Age, BMI, sex, family size, and educational level emerged as the most important predictors among the variables considered. - Authors conclude hypertension is common in Bangladesh with socio-demographic determinants; RF may be preferable for screening/public health applications because of higher recall, but external validation and clinical utility assessment are needed before implementation. - Data source is publicly available BDHS 2022 via DHS Program (requires registration). The study reported no specific funding and no competing interests.
## Clinical Analysis & Structured Key Points
Machine learning and deep learning–based prediction of hypertension and analysis of its major risk factors in Bangladesh | PLOS One Browse Subject Areas ? Click through the PLOS taxonomy to find articles in your field. For more information about PLOS Subject Areas, click here . Article Authors Metrics Comments Media Coverage Peer Review Reader Comments Figures Figures Abstract Background Hypertension is a leading cause of cardiovascular morbidity and mortality in Bangladesh. This study examined its prevalence, risk factors, and predictive modeling using machine learning (ML) and deep learning (DL) approaches. Method We analyzed cross-sectional data from the 2022 Bangladesh Demographic and Health Survey, which included 14,283 adults (≥18 years). Prevalence was estimated, chi-square tests assessed associations, and four ML models (weighted logistic regression, random forest, extreme gradient boosting, light gradient boosting machine) and two DL models (TabNet, and multi-layer perceptron) were applied to predict hypertension risk. Model performance was evaluated using accuracy, precision, recall, specificity, F1 score, and area under the receiver operating characteristics curve and precision-recall curve. Results Overall prevalence was 18.04% (95% CI: 17.2%–18.9%), higher among women (18.87%) than men (16.97%). The chi-square test suggests that hypertension was significantly associated with age, BMI, diabetes, wealth index, education, household size, and region (p < 0.05). Among the machine learning and deep learning models, weighted logistic regression (WLR) achieved the highest accuracy (0.817), precision (0.444), specificity (0.981), AUC-ROC (0.751), and AUC-PR (0.357). However, WLR exhibited low recall (0.070). In contrast, the random forest (RF) model achieved the highest recall (0.687) and F1-score (0.460) on the test data, indicating greater sensitivity in identifying individuals with hypertension. Additionally, age, BMI, sex, family size, and educational level were identified as the most important predictors among the variables included in the study. Conclusion Hypertension is common in Bangladesh, with higher prevalence in women and significant association with socio-demographic determinants. Although WLR demonstrated the highest accuracy, precision, specificity, and AUC-PR, its low recall limits its utility for identifying individuals with hypertension. RF may be more suitable for public health applications because of its higher recall and F1-score; however, further external validation and assessment of its clinical utility are required before implementation. Citation: Chandra S, Molla MM, Islam S, Rahaman MM, Ali M, Ali MA (2026) Machine learning and deep learning–based prediction of hypertension and analysis of its major risk factors in Bangladesh. PLoS One 21(9): e0358471. https://doi.org/10.1371/journal.pone.0358471 Editor: André Luis C. Ramalho, University of Porto Faculty of Medicine: Universidade do Porto Faculdade de Medicina, PORTUGAL Received: October 16, 2025; Accepted: September 1, 2026; Published: September 17, 2026 Copyright: © 2026 Chandra et al. This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Data Availability: The data underlying the results presented in this study on machine learning and deep learning–based prediction of hypertension and its major risk factors in Bangladesh were obtained from the Bangladesh Demographic and Health Survey (BDHS), available through the Demographic and Health Surveys (DHS) Program repository ( https://dhsprogram.com/data/available-datasets.cfm ). The data are publicly available; however, access requires registration and approval from the DHS Program. Due to data use restrictions, the authors are not permitted to share the data directly. The authors confirm that they had no special access privileges. Funding: The author(s) received no specific funding for this work. Competing interests: The authors have declared that no competing interests exist. Introduction Hypertension, commonly referred to as high blood pressure, it is a physical condition when the systolic blood pressure (SBP) ≥ 140 mmHg and/or diastolic blood pressure (DBP) ≥ 90 mmHg [ 1 ]. It is a major modifiable risk factor for cardiovascular diseases (CVDs), the leading cause of premature death [ 2 ]. Hypertension significantly increases the risk of various diseases [ 3 ], including heart attack [ 4 ], heart failure [ 5 ], coronary heart disease (CHD) [ 6 ], stroke [ 7 ], chronic kidney disease (CKD) [ 8 ], and diabetes [ 9 ]. Globally, hypertension affects an estimated 1.28 billion adults aged 30–79 years, two-thirds of whom live in LMICs [ 10 ]. By 2025, this figure is projected to reach 1.56 billion, driven by demographic changes, urbanization, dietary transitions, and sedentary lifestyles [ 11 – 13 ]. In 2019, CVDs caused about 18.6 million deaths, with nearly 80% occurring in low and middle-income countries (LMICs) [ 14 ]. While hypertension is a major global health concern, its impact is particularly alarming in Bangladesh, where 22.8 million adults aged 30–79 years have high blood pressure and more than 84% do not have their condition under control [ 15 ]. In Bangladesh, 14 of the 20 deaths were due to noncommunicable diseases, and the three leading causes were stroke, ischaemic heart disease, and chronic obstructive pulmonary disease [ 16 ]. According to the World Health Organization, approximately 283,000 people die from cardiovascular diseases in Bangladesh each year, of which 52% are linked to hypertension [ 15 ]. Furthermore, the expected annual cost of the hypertension control program was $3.2 million USD in Bangladesh [ 17 ]. In addition, increased life expectancy from improved healthcare has contributed to a growing burden of non-communicable diseases (NCDs) such as hypertension, whose prevalence is rising due to demographic, socioeconomic, and lifestyle factors [ 11 , 18 , 19 ]. Sex differences are evident, with men more prone earlier in life while postmenopausal women at greater risk due to reduced estrogen’s vasoprotective effects [ 20 ]. Metabolic factors, including elevated blood glucose and diabetes, share pathophysiological mechanisms with hypertension such as insulin resistance and vascular remodeling [ 21 ]. Socioeconomic status influences risk through healthcare access, diet quality, and stress exposure [ 22 ]. Environmental factors, including urban versus rural residence and regional differences, further influence hypertension prevalence; urbanization is associated with lifestyle patterns that promote hypertension [ 23 ]. Hypertension has been the focus of numerous studies [ 24 , 25 ], with a number of investigations conducted in Bangladesh using BDHS data [ 26 – 28 ]. For instance, Chowdhury et al., examined prevalence and risk factors among adults aged ≥35 years using logistic regression [ 29 ]. Iqbal et al., analyzed demographic, socioeconomic, and biological correlates, along with awareness, treatment, and control, using multivariate logistic regression [ 18 ]. Chowdhury et al., compared BDHS 2011 and 2017–18 data to assess changes in prevalence and risk factors, employing logistic regression and Wagstaff decomposition [ 19 ]. Sathi et al., explored trends in hypertension, diabetes, and their coexistence using modified Poisson regression [ 28 ]. Ghosh et al., estimated age-standardized prevalence and assessed determinants via hierarchical mixed-effects sequential Poisson regression [ 30 ]. Islam et al., applied feature selection (LASSO, SVM-RFE) and machine learning models (ANN, DT, RF, GB) to predict hypertension, with SVM-RFE combined with GB achieving the best performance [ 31 ]. Although machine learning methods are commonly used for hypertension prediction, applications of deep learning remain limited. In recent years, however, deep learning algorithms have been widely applied to predict a variety of diseases and have often outperformed traditional approaches. In this research, we employ both machine learning and deep learning techniques to predict hypertension using BDHS data, with the objective of identifying the most effective predictive model and the key risk factors associated with hypertension. We also investigated the prevalence of hypertension among adults’ people in Bangladesh. Materials and methods Study design, setting, participants, and study size Data for this study were obtained from the cross-sectional study of the Bangladesh Demographic and Health Survey (BDHS) 2022 [ 32 ]. Primary data were collected from June 27, 2022, to December 12, 2022, through face-to-face interviews, and biomarker measurements were taken by trained enumerators. BDHS 2022 was the third survey to include biomarker variables. It employed a two-stage stratified sampling design based on the integrated multi-purpose sampling master frame from a complete list of enumeration areas (EAs) covering the whole country. In the first stage, 674 EAs were selected using probability proportional to size; in the second stage, 45 households were randomly chosen per EA. Among 45 households, 30 households received the long questionnaire, where 15 were systematically selected for biomarker measurements, especially the women age between 15 and 49 and children under 5. Within this subsample, half of the households (8 of 15) were systematically selected for biomarker measurements among men aged ≥18 years, ever-married women aged ≥50 years, and never-married women aged ≥18 years, and blood pressure and blood glucose were measured for all men and women aged ≥18 years. Fig 1 shows that the 2022 BDHS included 30,330 households from 674 clusters, of which 20,220 were selected for the long women’s questionnaire and 5,392 for biomarker measurements. A total of 14,283 complete samples were obtained from 5,392 households, with no missing values across study variables. Download: PNG larger image TIFF original image Fig 1. BDHS 2022 sample selection procedure. https://doi.org/10.1371/journal.pone.0358471.g001 Ethical approval and consent to participates In this study, we used publicly available secondary data from the Bangladesh Demographic and Health Survey (BDHS). The survey was conducted under the authority of the National Institute of Population Research and Training and the Medical Education and Family Welfare Division, Ministry of Health and Family Welfare. The procedures and questionnaires for the BDHS survey are approved by the ICF institutional review board (IRB) and Bangladesh IBR. In order to collect the information, an informed consent statement was given to the respondent, and the participation in the survey was voluntary. Also, the identification numbers of the respondents, such as the enumeration area, household number, and individual number, are destroyed and randomly reassigned. As a result, individuals or households are not identifiable. Response variable and explanatory variables The variable of interest or response variable in this study is hypertension, commonly referred to as high blood pressure. Out of 30,330 households, 5,392 households were chosen as a measure of the biomarkers, and this measure includes the blood pressure. In which 6,853 men and 8,156 women aged 18 years or more could certainly undergo blood pressure and blood glucose level checkup. Blood pressure was measured in ninety-five percent of the female and ninety one percent of the male qualifiers. The respondents were measured three times each on blood pressure and the mean of the second and the third reading was taken. Participants were considered hypertensive if they SBP ≥ 140 mmHg or DBP ≥ 90 mmHg or were taking antihypertensive medication at that time to control it. We selected 10 explanatory variables, including age, sex, marital status, family size, division, residence type, wealth index, education level, body mass index (BMI), and diabetes status, which have also been used in previous studies [ 11 , 18 , 28 , 29 , 31 ]. The list of selected explanatory variables and their categories is provided in Table 1 . The BMI variable categorizes the participants as underweight (BMI < 18.5 kg/m 2 ), normal (18.5 ≤ BMI < 25 kg/ m 2 ) and obese/overweight (BMI ≥ 25 kg/ m 2 ) group. Diabetes status was determined based on plasma blood glucose levels (mmol/L); a person was considered diabetic if their glucose level exceeded 7 mmol/L or if they were taking medication for the condition, otherwise they were classified as non-diabetic [ 33 ]. The variable family size is determined by the members of the family: small family has members less than 5, medium family has 5–7 members, and large family has more than 7 members. Download: PNG larger image TIFF original image Table 1. Description of the response and explanatory variable(s). https://doi.org/10.1371/journal.pone.0358471.t001 Statistical analysis The baseline characteristics of the survey are indicated by frequency (%), at each variable. Also, the prevalence of hypertension is calculated by category with the help of the confidence interval (CI). Multicollinearity was assessed using the variance inflation factor (VIF). The value less than 5 indicates no major multicollinearity. The chi-square test was applied in determining the significance of the variables and 0.05 value of the significance level would be considered. Additionally, six machine learning and deep learning models were used to identify the risk factors of hypertension. For statistical analysis, we use R version 4.5.1. To incorporate the complex survey design, we use the survey package and svydesign functions in R programming. Dataset pre-processing Handling the class imbalance of data. Class imbalance occurs when samples of one class are significantly higher than the other class. A typical dataset is considered imbalanced if the majority class observations are twice or more than the minority class [ 34 ]. In our study, the response variable, hypertension, was divided into two categories: yes and no. The minority class (yes) consists of 18.04% of the data, while the majority class (no) consists of 81.96% of the data. Therefore, our dataset is imbalanced. We apply the synthetic minority oversampling technique (SMOTE) to address the class imbalance. SMOTE is applied only for the training data, not for test data. SMOTE was used for all the models except the weighted logistic regression model. Missing data and categorical variables. No imputation technique was used to handle missing data; we simply deleted the entire row from the dataset. The hypertension variable was encoded as a binary class (Yes = 1, No = 0), while other categorical predictors (sex, age group, family size, marital status, division, residence, wealth index, education, BMI, and diabetes) were encoded using one-hot encoding. To optimize scale-sensitive models such as MLP, features were standardized to zero mean and unit variance using StandardScaler, whereas no scaling was applied to other models (Weighted Logistic Regression, Random Forest, XGBoost, LightGBM, TabNet). Train-test split. We divided our dataset into two portion that is training and test data. For the training data, we kept 80% of it, and for the test, we kept the remaining 20%. During the train test split we stratified the data by the hypertension status. The 80/20 split was not explicitly constrained by survey cluster, observations from the same cluster could potentially occur in both the training and test sets. To assess the potential influence of cluster-level dependence, we additionally performed 5-fold and 10-fold StratifiedGroupKFold cross-validation, using survey cluster as the grouping variable. ML and DL models and parameter tuning In our study we have used four machine learning (ML) and two deep learning (DL) models, including weighted logistic regression, random forest, eXtreme Gradient Boosting, Light Gradient Boosting Machine, TabNet, and multi-layer perceptron (MLP). In this study, the DHS sampling weights were used only in the weighted logistic regression, which is the baseline survey-weighted regression model. DHS sampling weights were not used in the training of the random forest, XGBoost, LightGBM, MLP, and TabNet models. We made this distinction because the purpose of the analysis is to compare various predictive performance of various ML and DL algorithms with a survey-weighted regression model. Weighted logistic regression. Let, there be a set of independent variables and a binary dependent variable y. Now we have a set of n independent observations, for where are the value of corresponding , is the corresponding weight, and is the value of observation. Then the logistic regression model is defined as follows: (1) Where, are the regression coefficient. Then, the likelihood function of weighted logistic regression is defined as follow: (2) The weighted logistic regression is useful when the data is imbalanced. It can handle the imbalanced data without synthetic data generation or resampling techniques [ 35 , 36 ]. Random forest (RF). Random Forest is an ensemble algorithm that sequentially builds a number of decision trees, and based on those trees makes a prediction more accurate and stable model [ 37 ]. For hyperparameters in the RF model, we use the GridSearchCV class in sklearn.model_selection module where in param_grid we use n_estimators = [20, 50, 100, 150, 200, 300] and obtained the best parameter n_estimators = 150, the other parameters were max_depth = 10, min_samples_leaf = 30, class_weight = “balanced”, and random_state = 42. Light Gradient Boosting Machine (LightGBM). LightGBM is an ensemble learning framework based on gradient boosting, which builds a strong predictive model by sequentially adding decision trees that minimize a loss function via gradient descent [ 38 ]. For hyperparameters in the LightGBM model, we use the GridSearchCV class in sklearn.model_selection module where in param_grid we use parameters num_leaves = [5, 20, 30, 50], learning_rate = [0.01, 0.05, 0.1, 0.2], and n_estimators: [20, 50, 100, 150, 200, 300] and obtained the best parameters learning_rate = 0.2 n_estimators = 200, and num_leaves = 30. Extreme gradient boosting (XGBoost). XGBoost is a powerful and fast implementation of the gradient boosting framework, optimized to work with less structured data (usually structured data such as a table), with incredible predictive performance. Theoretically, it is well suited to machine learning competitions and practical applications due to its robustness, regularization properties, and computational efficiency [ 39 ]. For hyperparameters in the LightGBM model, we use the GridSearchCV class in sklearn.model_selection module where in param_grid we use parameters max_depth = [3, 5, 7, 9], learning_rate = [0.01, 0.05, 0.1], n_estimators = [25, 50, 100, 200, 300], subsample = [0.5,0.7, 0.8, 0.9], and colsample_bytree = [0.7, 0.8, 0.9] and obtained the best parameters max_depth = 9, learning_rate = 0.05, n_estimators = 300, subsample = 0.8, and colsample_bytree = 0.7. TabNet. TabNet is a deep learning model that is tailor-made to tabular data to unite the capacity of neural networks with interpretability in the attention mechanism. In contrast to traditional neural networks which treat features differently, TabNet utilizes sequential attention to make decisions by attending to most relevant ones at each decision step, enhancing performance and providing explainability [ 40 ]. To define the TabNet model we use parameter n_d = 16, n_a = 16, n_steps = 4, gamma = 1.5, lambda_sparse = 1e-3, optimizer_params = dict(lr = 2e-2), and verbose = 1. During the fitting of the model, we use max_epochs = 100, patience = 20, and batch_size = 32. Multi-layer perceptron (MLP). Multi-Layer Perceptron (MLP) is a form of feedforward artificial Neural network which i
## Related Clinical Research

- [Lifestyle Changes May Help Prevent Multimorbidity In Prediabetes](https://medichelpline.com/clinical-feed/medical-news-today-23-diet-weight-loss-and-150-minutes-of-exercise-may-protect-against-chronic.md)
- [Moving Forward: The Future of GLP-1 Therapies](https://medichelpline.com/clinical-feed/endocrine-news-18-moving-forward-the-future-of-glp-1-therapies.md)
- [Multimodal extraction of heart disease risk factors from free-text and structured records using Li](https://medichelpline.com/clinical-feed/plos-one-0-leveraging-free-text-clinical-records-for-heart-disease-classification-through.md)
- [Caffeinated Coffee and Heart Health: How Much Coffee Is Safe Daily?](https://medichelpline.com/clinical-feed/aha-news-0-coffee-and-heart-health-how-many-cups-of-caffeinated-coffee-are-safe-to-drink.md)
- [Cardiovascular and renal outcomes of combined SGLT2 inhibitors and GLP-1 receptor agonists versus monotherapy in patients with type 2 diabetes mellitus: a network meta-analysis [Research]](https://medichelpline.com/clinical-feed/cmaj-0-cardiovascular-and-renal-outcomes-of-combined-sglt2-inhibitors-and-glp-1.md)

## Navigation
- [← Back to Cardiology Feed](https://medichelpline.com/clinical-feed/cardiology.md)
- [← All Clinical Specialties](https://medichelpline.com/clinical-feed.md)
## Medical & Regulatory Disclaimer

> [!CAUTION]
> MedicHelpline content is structured for research, educational, and professional discovery purposes. It does not constitute individual medical advice, clinical diagnosis, or treatment recommendations.
> Always verify dosing, contraindications, and regulatory alerts against official product labeling and primary regulatory sources before clinical decision-making.