This retrospective multicenter study evaluated whether combining electrocardiography (ECG) and chest radiography (CXR) in a single deep learning model improves prediction of incident moderate-to-severe regurgitant valvular heart disease (rVHD). The cohort comprised 212,888 paired ECG–CXR examinations from 116,380 patients collected at two Chinese centers. Baseline ECG and CXR were required to be obtained within 60 days of an echocardiogram, which served as the diagnostic reference and provided outcome ascertainment for progression to moderate-to-severe aortic regurgitation (AR), mitral regurgitation (MR), or tricuspid regurgitation (TR).
The modeling framework used pretrained unimodal encoders for ECG and CXR, a token-level cross-modal fusion mechanism to integrate features, and a class-specific gating module that adaptively weighted ECG-only, CXR-only, and fused predictions. The approach therefore aimed to leverage complementary electrical signals from ECG and structural/hemodynamic information from CXR in a single predictive pipeline.
Model performance was assessed with time-to-event and classification-oriented metrics, including concordance index (C-index), area under the receiver-operating-characteristic curve (AUROC), area under the precision-recall curve (AUPRC), decision curve analysis for clinical net benefit, net reclassification improvement (NRI) across multiple time horizons (1–5 years), and Kaplan–Meier stratification for risk groups. The authors compared the multimodal fused model against ECG-only and CXR-only unimodal models across all three valve phenotypes.
For AR, the multimodal fused model showed the largest relative improvement over the ECG-only model. ECG-only discrimination for AR had a C-index of 0.616. The multimodal model increased C-index to 0.713. Reported classification metrics for AR included AUROC of 0.729 and AUPRC of 0.972 for the multimodal model. These results indicate that adding CXR-derived information to ECG predictions materially improved risk discrimination for future moderate-to-severe AR in this cohort.
For MR, multimodal fusion also improved performance relative to unimodal inputs. The multimodal C-index was 0.801, with AUROC 0.814 and AUPRC 0.972. By comparison, ECG-only and CXR-only models had C-indices of 0.782 and 0.775, respectively. Thus, the fused model achieved better discrimination than either modality alone for incident moderate-to-severe MR.
For TR, the multimodal model and the CXR-only model achieved similar discrimination by C-index (both reported at 0.802). However, decision curve analysis showed that multimodal fusion offered greater net benefit across clinically relevant risk thresholds, suggesting improved clinical utility despite similar discrimination metrics. NRI analyses were positive for TR across evaluated time horizons, indicating improved reclassification with multimodal predictions.
The authors applied Grad‑CAM–style interpretability analyses to examine which ECG leads and which CXR regions the models emphasized. ECG attention maps localized to leads consistent with expected electrical signatures: leads II and precordial leads V4–V6 were highlighted for AR; leads I, II, aVF, and V4–V6 for MR; and inferior and right precordial leads for TR. CXR attention emphasized patterns such as chamber-specific enlargement and pulmonary congestion that align with known structural and hemodynamic manifestations of regurgitant lesions. The interpretability results were presented as evidence that the multimodal model leveraged complementary, biologically plausible signals from both modalities.
The authors note that both ECG and CXR are widely available, low-cost tests in routine care, which supports scalability of a multimodal screening or risk‑stratification approach. Quantitatively, NRI was positive across 1–5 year horizons for each valve type, and decision curve analysis indicated net clinical benefit for multimodal fusion in multiple scenarios.
Limitations reported in the source include the retrospective design and restricted access to underlying data due to patient privacy and institutional regulations. The report is a preprint and has not been peer reviewed; the authors explicitly recommend prospective studies to validate clinical implementation. Details on external validation cohorts, calibration metrics beyond C-index/AUROC/AUPRC, and any deployment considerations (hardware, inference latency, or integration with clinical workflows) were not reported in the source abstract.
In this large retrospective cohort, a multimodal deep learning model that integrates ECG and CXR improved prediction of incident moderate-to-severe regurgitant valvular heart disease compared with models using a single modality. The greatest incremental benefit was observed for aortic regurgitation. Interpretability analyses supported biological plausibility by showing modality-specific attention patterns consistent with known pathophysiology. Given the study’s retrospective nature, restricted data availability, and preprint status, the authors recommend prospective validation before clinical adoption. If validated, the approach could offer a scalable risk-stratification tool that capitalizes on routinely acquired, low‑cost tests to prioritize echocardiographic evaluation.