---
title: "Multimodal extraction of heart disease risk factors from free-text and structured records using Li"
id: "plos-one-0-leveraging-free-text-clinical-records-for-heart-disease-classification-through"
canonical_url: "https://medichelpline.com/clinical-feed/plos-one-0-leveraging-free-text-clinical-records-for-heart-disease-classification-through"
content_type: "clinical_feed_article"
specialty: "Cardiology"
source_name: "PLOS ONE (Medicine)"
source_url: "https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916"
published_at: "2026-07-27T14:00:00.000Z"
evidence_level: "Journal Feed"
license: "CC-BY-NC-4.0 / Informational Use"
---
# Multimodal extraction of heart disease risk factors from free-text and structured records using Li
## Provenance & Clinical Metadata
- **Canonical URL:** https://medichelpline.com/clinical-feed/plos-one-0-leveraging-free-text-clinical-records-for-heart-disease-classification-through
- **Specialty:** [Cardiology](https://medichelpline.com/clinical-feed/cardiology.md)
- **Primary Source:** PLOS ONE (Medicine)
- **Source URL:** [Original Journal Publication](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916)
- **Published At:** 2026-07-27T14:00:00.000Z
- **Evidence Rating:** Journal Feed
## Executive GIST (TL;DR)
- Heart disease remains a leading cause of global mortality; **hypertension** and **diabetes** are major, often co-occurring risk factors that accelerate coronary artery disease and related complications. - Clinical risk-factor information is frequently stored as **unstructured free text** in EHRs, complicating retrieval and automated analysis with traditional methods. - This study integrated two complementary data sources: the PrevComp corpus (274 deidentified EHRs from patients with both hypertension and diabetes) and the UCI Heart Disease dataset (303 patients referred for coronary angiography) to build a multimodal classification pipeline. - The UCI dataset provided 13 structured features (age, sex, chest pain type, resting BP, serum cholesterol, fasting blood sugar, resting ECG results, max heart rate, exercise-induced angina, ST depression/old peak, slope of ST segment, number of major vessels, thalassemia) and was split 60/20/20 for training/validation/test with exhaustive GridSearchCV (3-fold stratified CV) for hyperparameter tuning. - Tree-based ensemble methods were compared (LightGBM, XGBoost, CatBoost, random forest); **LightGBM** (the UCI-trained LightGBM model) produced the best performance in this setting. - To convert free-text notes into structured-like features compatible with the UCI model, each PrevComp note was tokenized and processed with **ClinicalBERT** to extract a 768-dimensional CLS embedding as a fixed-length representation. - Because ground-truth structured labels were not available in PrevComp, a large language model (Meta-Llama-3.1) was used to generate weakly supervised labels for each note, enabling downstream evaluation with the UCI-trained model. - The multimodal classification pipeline achieved an overall predictive accuracy of 83% for presence/absence of heart disease, supporting the feasibility of mapping unstructured clinical narratives to structured risk-factor representations and classifying heart disease status. - Experiments were implemented in Python on Google Colab with 25 GB of RAM, and all relevant data sources are available on Kaggle as reported by the authors. - The work is reported as the first use of the PrevComp corpus for classifying records by the presence/absence of heart disease and demonstrates the potential of combining **NLP-derived features** and structured-model training for clinical risk detection.
## Clinical Analysis & Structured Key Points
[ Skip to main content ](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#main-content) Advertisement * [plos.org](https://plos.org/) * [Create account](https://community.plos.org/registration/new) * [Sign in](https://journals.plos.org/user/secure/login?page=%2Fplosone%2Farticle%3Fid%3D10.1371%2Fjournal.pone.0354916) * * About * Browse * Publish * [](https://journals.plos.org/plosone/ "PLOS One") * Search [advanced search](https://journals.plos.org/plosone/search) * [Browse Topics](https://journals.plos.org/plosone/subjectAreaBrowse) Browse Subject Areas ? Click through the PLOS taxonomy to find articles in your field. For more information about PLOS Subject Areas, click [here](https://github.com/PLOS/plos-thesaurus/blob/master/README.md "Link opens in new window"). [](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916) [](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916) * 0 [Save](https://journals.plos.org/plosone/article/metrics?id=10.1371/journal.pone.0354916#savedHeader) [Total Mendeley and Citeulike bookmarks.](https://journals.plos.org/plosone/article/metrics?id=10.1371/journal.pone.0354916#savedHeader) * 0 [Citation](https://journals.plos.org/plosone/article/metrics?id=10.1371/journal.pone.0354916#citedHeader) [Paper's citation count computed by Dimensions.](https://journals.plos.org/plosone/article/metrics?id=10.1371/journal.pone.0354916#citedHeader) * 28 [View](https://journals.plos.org/plosone/article/metrics?id=10.1371/journal.pone.0354916#viewedHeader) [PLOS views and downloads.](https://journals.plos.org/plosone/article/metrics?id=10.1371/journal.pone.0354916#viewedHeader) * 0 [Share](https://journals.plos.org/plosone/article/metrics?id=10.1371/journal.pone.0354916#discussedHeader) [Sum of Facebook, Twitter, Reddit and Wikipedia activity.](https://journals.plos.org/plosone/article/metrics?id=10.1371/journal.pone.0354916#discussedHeader) Open Access Peer-reviewed Research Article # Leveraging free-text clinical records for heart disease classification through structured feature mapping * Noha Alnazzawi Roles Data curation, Funding acquisition, Project administration, Resources, Supervision, Writing – original draft, Writing – review & editing * E-mail: alnazzawin@rcjy.edu.sa Affiliation Department of Computer Science and Engineering, Yanbu Industrial College, Royal Commission for Jubail and Yanbu, Yanbu, Saudi Arabia [ ![ORCID logo](https://journals.plos.org/resource/img/orcid_16x16.png) https://orcid.org/0000-0003-3265-0297 ](https://orcid.org/0000-0003-3265-0297 "ORCID Registry") ⨯ # Leveraging free-text clinical records for heart disease classification through structured feature mapping * Noha Alnazzawi ![PLOS](https://journals.plos.org/resource/img/logo-plos-full-color.svg) x * Published: July 27, 2026 * * [Article](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916) * [Authors](https://journals.plos.org/plosone/article/authors?id=10.1371/journal.pone.0354916) * [Metrics](https://journals.plos.org/plosone/article/metrics?id=10.1371/journal.pone.0354916) * [Comments](https://journals.plos.org/plosone/article/comments?id=10.1371/journal.pone.0354916) * [Media Coverage](http://plos.altmetric.com/details/doi/10.1371/journal.pone.0354916) * [Abstract](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#abstract0) * [Introduction](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#sec001) * [Related work](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#sec002) * [Materials and methods](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#sec003) * [Results and discussion](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#sec007) * [Conclusions](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#sec008) * [References](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#references) * [Reader Comments](https://journals.plos.org/plosone/article/comments?id=10.1371/journal.pone.0354916) * [Figures](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916) ## Abstract Hypertension and diabetes are major risk factors for heart disease, which remains among the leading causes of morbidity and mortality worldwide. Heart disease includes heart failure, myocardial infarction, stroke, and atherosclerosis. The identification and monitoring of these risk factors are crucial for early intervention and effective management. Machine learning techniques have the potential to improve the management and prevention of heart disease by enabling the automatic detection of risk factors, which in turn can help doctors personalize treatment and facilitate preventive interventions. In this study, heart disease risk factors were automatically extracted using a combination of multimodal data: both unstructured clinical narratives (e.g., the PrevComp corpus) and structured datasets (e.g., the UCI heart disease dataset) were used to predict the presence or absence of heart disease. The classification model is based on the Light Gradient Boosting Machine (LightGBM), a state-of-the-art implementation of the gradient boosting framework that employs tree-based learning algorithms. The developed classification model demonstrated a robust predictive accuracy of 83% for the presence/absence of heart disease, supporting its potential to accurately identify high-risk patients and improve clinical outcomes. ## Figures ![Table 3](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0354916.t003) ![Table 4](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0354916.t004) ![Table 5](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0354916.t005) ![Fig 1](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0354916.g001) ![Table 1](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0354916.t001) ![Table 2](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0354916.t002) ![Table 3](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0354916.t003) ![Table 4](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0354916.t004) ![Table 5](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0354916.t005) ![Fig 1](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0354916.g001) ![Table 1](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0354916.t001) ![Table 2](https://journals.plos.org/plosone/article/figure/image?size=inline&id=10.1371/journal.pone.0354916.t002) **Citation:** Alnazzawi N (2026) Leveraging free-text clinical records for heart disease classification through structured feature mapping. PLoS One 21(7): e0354916. https://doi.org/10.1371/journal.pone.0354916 **Editor:** Agnese Sbrollini, Polytechnic University of Marche: Universita Politecnica delle Marche, ITALY **Received:** July 11, 2025; **Accepted:** July 14, 2026; **Published:** July 27, 2026 **Copyright:** © 2026 Noha Alnazzawi. This is an open access article distributed under the terms of the [Creative Commons Attribution License](http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. **Data Availability:** All relevant data are available on Kaggle at and . **Funding:** The author(s) received no specific funding for this work. **Competing interests:** The authors have declared that no competing interests exist. ## Introduction According to the latest World Health Organization (WHO) report [[1](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref001)], more than 17.9 million people were estimated to have died in 2019 because of cardiovascular disease, accounting for 32% of all deaths globally. Among these deaths, 85% were due to heart attack or stroke. Owing to its diverse complexity and extremely high mortality rates, heart disease has become a formidable medical challenge [[2](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref002)]. The main risk factors that contribute to coronary artery disease (CAD) are sex, smoking, older age, family history, poor diet, lipid levels, lack of physical activity, hypertension, weight gain, and alcohol consumption [[3](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref003)]. Hypertension and diabetes are two examples of risk factors that can be inherited and increase the likelihood of developing heart disease [[4](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref004)]. Up to 75% of adults with diabetes also have hypertension, and patients with hypertension alone often show evidence of insulin resistance. Thus, hypertension and diabetes are common conditions that significantly overlap in terms of underlying risk factors and associated complications [[5](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref005)]. Hypertension increases the load on the heart, leading to heart muscle thinning over time. It also causes damage to the walls of blood vessels, which accelerates the development of atherosclerosis, the major cause of CAD. Over time, hypertension causes damage to the arteries, which in turn increases the risk of blood clots, strokes, and heart attacks. On the other hand, diabetes leads to the accumulation of glucose in the blood, which can damage blood vessels, accelerate atherosclerosis, and trigger inflammation in blood vessels, further promoting the buildup of fatty plaques that narrow and harden the arteries [[6](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref006),[7](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref007)]. When present together, hypertension and diabetes significantly increase the risk of heart disease and can accelerate the progression of CAD. Notably, understanding the risk factors for heart disease in patients with hypertension and diabetes can not only guide preventive and risk-reduction measures but can also help manage patients with existing heart disease. The automatic detection of heart disease in patients with hypertension and diabetes will help doctors effectively screen for risk factors, identify patients at increased risk and facilitate early interventions for these diseases. Furthermore, automatic detection can aid the customization of treatment plans for individuals on the basis of their specific risk profiles [[6](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref006),[8](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref008)]. However, information related to heart disease risk factors is often recorded in an unstructured format, namely free text, in electronic health records (EHRs), making data retrieval difficult and inefficient with traditional data processing methods. Advances in text mining (TM) techniques provide efficient means to automate the extraction and integration of vital information, including disease risk factors and complications, from EHRs, allowing for more effective handling of large volumes of unstructured clinical data. The PrevComp corpus [[9](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref009)] consists of text obtained from the EHRs of patients known to have both hypertension and diabetes, including disease complications. For the PrevComp corpus, different machine learning techniques are used to automate the extraction and integration of disease complications from narrative text. In contrast, the Cleveland Heart Disease dataset from the University of California Irvine (UCI) Machine Learning Repository contains structured data on heart disease risk factors. It is typically used to train machine learning models to predict the presence or absence of heart disease on the basis of various risk factors [[10](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref010)]. The UCI dataset includes various features that are relevant for diagnosing heart disease, such as age, sex, chest pain type, resting blood pressure, cholesterol level, fasting blood sugar, resting electrocardiographic results, maximum heart rate, exercise-induced angina, ST depression induced by exercise relative to rest, slope of the peak exercise ST segment, number of major vessels (0–3) colored by fluoroscopy, and thalassemia. The aim of this research is to utilize advanced machine learning (ML) techniques to train an ML model on integrated multimodal data by combining structured data (e.g., the ICU dataset) with unstructured data (e.g., clinical notes from the PrevComp corpus) to provide a more comprehensive understanding of a patient’s condition. By combining these diverse data types, machine learning models can assist doctors in making more accurate and informed decisions for diagnosis, treatment, and personalized care. Although several studies have identified the risk factors for heart disease, to our knowledge, our work represents the first attempt at integrating multimodal data (i.e., structured and unstructured) to classify narrative clinical records according to the presence or absence of heart disease as a complication of hypertension and diabetes. The contributions of this article are twofold: 1. This is the first study in which the PrevComp corpus is used to classify patient records on the basis of the presence or absence of heart disease in the patient. 2. Different ML models were trained on multimodal data via structured data (i.e., the UCI dataset) to extract heart disease risk factors from unstructured data (i.e., clinical notes from the PrevComp corpus) and, on the basis of risk factors, classify clinical records according to the presence or absence of heart disease in the patient. ## Related work Text mining in health care is crucial because much of the relevant clinical information exists in EHRs. TM has shown significant promise in extracting heart disease risk factors from clinical records, offering the potential for better disease management, early detection, and personalized treatment [[11](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref011)]. In a prospective cohort study, Weng et al. [[12](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref012)] demonstrated that machine learning algorithms significantly enhanced the prediction of cardiovascular disease, highlighting their effectiveness and practical applicability in this domain. In the health care domain, machine learning is particularly valuable for predicting patient outcomes, such as hospital readmissions, when large, complex datasets are analyzed. This facilitates timely interventions and enhances care delivery. ML techniques have already been successfully applied to predict a variety of clinical events, such as cardiovascular events [[13](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref013)], sepsis [[14](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref014)], delirium [[15](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref015)], and hospital readmissions following lumbar laminectomy [[16](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref016),[17](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref017)]. EHRs consist of two types of data, namely, structured and unstructured data—both of which provide complementary information. Structured data represent discrete, organized and predefined data, such as demographic data, laboratory results, medication lists, and diagnosis codes, whereas unstructured data provide narrative details, such as family history, chief complaints, signs and symptoms [[18](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref018),[19](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref019)]. On the one hand, an abundance of research has been conducted on unimodal data using either structured or unstructured clinical data for heart disease prediction using different ensemble machine learning algorithms and deep learning, which have consistently shown high accuracy when different feature selection and balancing techniques are used [[20](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref020)–[24](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref024)]. For example, when structured EHR data are used, gradient boosting decision trees (GBDTs) have been widely recognized for their high accuracy and robustness in medical classification tasks, such as the prediction of heart disease. GBDT models such as extreme gradient boosting (XGBoost) [[25](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref025)], light gradient boosting machine (LightGBM) [[26](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref026)], and categorical boosting (CatBoost) [[27](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref027)] have demonstrated superior performance on various clinical datasets because of their ability to handle feature interactions, missing values, and heterogeneous data sources [[28](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref028),[29](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref029)]. LightGBM [[26](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref026)] is a high-performance machine learning algorithm that is widely used for classification, regression, and ranking tasks [[30](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref030)]. As a member of the gradient boosting family, LightGBM is known for its computational efficiency, high accuracy, and low memory usage. The LightGBM algorithm achieves exceptional performance in a variety of applications, such as multiclass classification [[31](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref031)], click prediction [[32](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref032)], and learning to rank tasks [[33](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref033)]. In several studies, the LightGBM algorithm has been successfully applied in both classification and regression problems and has consistently revealed excellent detection results, highlighting its effectiveness as a predictive model [[30](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0354916#pone.0354916.ref030)]. In comparative evaluations, the LightGBM algorithm outperforms other advanced machine learning methods in various classification and diagnos
## Related Clinical Research

- [Lifestyle Changes May Help Prevent Multimorbidity In Prediabetes](https://medichelpline.com/clinical-feed/medical-news-today-23-diet-weight-loss-and-150-minutes-of-exercise-may-protect-against-chronic.md)
- [Genistein modifies the association between EASIX and estimated 10-year ASCVD risk](https://medichelpline.com/clinical-feed/plos-one-13-interaction-between-genistein-and-the-endothelial-activation-and-stress-index.md)
- [Caffeinated Coffee and Heart Health: How Much Coffee Is Safe Daily?](https://medichelpline.com/clinical-feed/aha-news-0-coffee-and-heart-health-how-many-cups-of-caffeinated-coffee-are-safe-to-drink.md)
- [Cardiovascular and renal outcomes of combined SGLT2 inhibitors and GLP-1 receptor agonists versus monotherapy in patients with type 2 diabetes mellitus: a network meta-analysis [Research]](https://medichelpline.com/clinical-feed/cmaj-0-cardiovascular-and-renal-outcomes-of-combined-sglt2-inhibitors-and-glp-1.md)
- [Moving Forward: The Future of GLP-1 Therapies](https://medichelpline.com/clinical-feed/endocrine-news-18-moving-forward-the-future-of-glp-1-therapies.md)

## Navigation
- [← Back to Cardiology Feed](https://medichelpline.com/clinical-feed/cardiology.md)
- [← All Clinical Specialties](https://medichelpline.com/clinical-feed.md)
## Medical & Regulatory Disclaimer

> [!CAUTION]
> MedicHelpline content is structured for research, educational, and professional discovery purposes. It does not constitute individual medical advice, clinical diagnosis, or treatment recommendations.
> Always verify dosing, contraindications, and regulatory alerts against official product labeling and primary regulatory sources before clinical decision-making.