AI learned from cardiologists’ notes, and needed far less labeled ECG data

By teaching an AI system to connect raw heart signals with the clinical insights contained in cardiologists’ reports, researchers tested whether routine ECGs could reveal more with far fewer labeled examples.

Study: Development and external validation of a contrastive learning foundation model for ECG-based prediction of cardiovascular diseases and outcomes. Image Credit: ArmadilloPhotograp / Shutterstock

In a recent study published in The Lancet Digital Health, researchers in the United States revealed the development and external validation of ECG-CLIP, a novel multi-modal artificial intelligence (AI) foundation model for electrocardiogram (ECG) analysis. The model was developed using a large Scripps Health dataset of raw 12-lead ECG signals paired with clinician-overread reports to improve its diagnostic generalizability.

The researchers found that ECG-CLIP consistently outperformed conventional supervised architectures and a non-ECG foundation model across diverse tasks, including acute myocardial infarction detection, cardiomyopathy screening, and future atrial fibrillation prediction. Compared with ECG-specific foundation models, its advantages were most pronounced when only a small amount of labeled training data was available.

The model was further shown to match the performance of next-best comparators trained on full datasets while requiring substantially less labeled data, with reductions ranging from 83.1% to 93.7% depending on the task, highlighting its benefits and potential applicability in modern clinical settings.

The findings highlight the potential of ECG-CLIP as a versatile framework for extracting rich physiological representations from cardiac waveforms. The authors posit that the model offers substantial potential for future use in resource-constrained settings and wearable platforms, although prospective, multisite validation is required before these applications can be established clinically.

“Our new algorithm only needs to see on the order of a dozen confirmed ECGs of a specific disease to detect that disease in the future,” says senior author Giorgio Quer, an assistant professor of digital medicine at Scripps Research. “This is similar to how a clinician would learn: not from a million examples, but from understanding the general physiology behind an ECG first and then seeing a few specific cases.”

Background

Cardiovascular disease (CVD) has long remained one of the foremost global causes of clinical morbidity and mortality, with current public health reports estimating that ~20% of all deaths in the United States (US) are associated with CVD. Early diagnosis is key to prompt intervention and can help reduce the risk of serious consequences.

Consequently, the 12-lead electrocardiogram (ECG) remains one of the most important diagnostic tools for mitigating the impact of CVD and guiding clinical intervention. More than 300 million ECGs are performed worldwide each year, and variations in practitioner expertise and limited time or visual acuity can make subtle but clinically meaningful ECG changes difficult to detect.

Differences in expertise, limited clinician time, and subtle ECG changes can therefore contribute to inconsistent interpretation, underscoring the need for reliable automation in the field. While advances in computational throughput and supervised deep learning algorithms have partially addressed this need and facilitated automated ECG interpretation, most extant models are reported to be rigidly bound to task-specific labels.

These early deep learning implementations have been found to exhibit performance declines when applied to data acquired using different equipment or at different facilities.

Furthermore, most ECG foundation models have relied predominantly on signal-only pretraining. Previous approaches incorporating text supervision have largely relied on machine-generated or structured report strings rather than clinician-overread reports, thereby limiting direct alignment between complex voltage signals and the nuanced diagnostic information embedded in natural-language clinical reports.

About the study

The present study aimed to address these empirical constraints and support future CVD diagnosis by developing ECG-CLIP, a contrastive learning foundation model that pairs diagnostic ECGs with clinician-overread reports to enhance its generalizability and clinical applicability compared with currently available ECG foundation models.

The study used data from the Scripps Health system (2008–2019), which comprised 1,679,884 diagnostic ECGs and paired clinician-overread reports from 538,028 patients after duplicate, missing, and corrupted records were excluded. Patients were divided into development and held-out evaluation partitions. Model pretraining employed a two-stage architecture in which ECG-CLIP was first trained on self-supervised masked ECG reconstruction, followed by multi-modal contrastive learning to align ECG patterns with information from paired clinical reports.

Following model training, external performance validation was conducted using data from the independent MIMIC-IV database (800,035 ECGs from 161,352 patients). These validations comprised three primary task categories: 1. Cardiovascular disease detection, 2. Atrial fibrillation prediction, and 3. Adverse health outcome prediction.

Model performance on MIMIC-IV tasks was evaluated using five-fold cross-validation across label-availability thresholds ranging from 10 to 1,000 positive labels. Additional Scripps-based analyses used separate training and validation sets, as well as a held-out test set, including evaluations of single-lead ECG configurations.

“We found that ECG-CLIP was better at detecting and predicting cardiovascular diseases, particularly in cases where there was much less data,” says Quer. “This may be particularly useful in cases like rare diseases, where there are only a dozen or so positive examples of well-labeled ECGs that can be used for training the model.”

Study findings

ECG-CLIP showed diagnostic performance that was significantly superior to both ResNet (a ‘baseline’ supervised model; P < 0.0001) and OpenAI CLIP (a general vision-language foundation model without ECG-specific pretraining; P < 0.0001). Specifically, ECG-CLIP achieved AUCs of 0.963 for acute myocardial infarction (AMI) detection, 0.852 for cardiac amyloidosis, and 0.861 for hypertrophic cardiomyopathy (HCM).

ECG-CLIP achieved an AUC of 0.867 for predicting future atrial fibrillation from ECGs recorded during normal sinus rhythm. When predicting future adverse events, the model achieved an AUC of 0.855 for 30-day emergency department mortality, 0.797 for 30-day post-surgical mortality, 0.802 for 3-year chronic kidney disease onset, and 0.729 for 3-year type 2 diabetes onset.

In the label-efficiency analyses, ECG-CLIP could match the performance of the next-best comparator trained on the full dataset while using 88.9% to 92.7% less data across the cardiovascular disease detection and atrial fibrillation tasks. Across adverse outcome tasks, the reduction ranged from 83.1% to 93.7%.

The authors noted that DeepECG-SSL included MIMIC-IV data in its pretraining corpus, which could have influenced direct performance comparisons on that dataset.

Finally, single-lead evaluations demonstrated ECG-CLIP’s strong lead efficiency. The model was shown to retain useful discrimination for anterior AMI using less informative leads, including leads I and II, with AUCs above 0.75, while OpenAI CLIP and ResNet performed substantially worse.

Saliency analyses also indicated that ECG-CLIP focused on clinically relevant ST-segment regions, including leads V2-V4 for anterior AMI and leads I and II for inferior AMI, supporting the physiological interpretability of its predictions.

Conclusions

The present study demonstrates the potential of ECG-CLIP as an adaptable foundation model capable of learning transferable representations from routine natural language reports and integrating these insights with electrophysiological signals, highlighting its advantages, particularly when labeled training data are scarce.

While the authors note that the study was limited by external MIMIC-IV ECG labels that were not reviewed by cardiologists, development and evaluation in controlled hospital environments, and potential bias or noise in institution-specific clinician reports, the results underscore ECG-CLIP’s potential for future clinical and wearable applications. However, validation across multiple sites, diverse populations, and recording systems, together with prospective clinical trials, is required before real-world deployment can be established.

Source:
Journal reference:
Hugo Francisco de Souza

Written by

Hugo Francisco de Souza

Hugo Francisco de Souza is a scientific writer based in Bangalore, Karnataka, India. His academic passions lie in biogeography, evolutionary biology, and herpetology. He is currently pursuing his Ph.D. from the Centre for Ecological Sciences, Indian Institute of Science, where he studies the origins, dispersal, and speciation of wetland-associated snakes. Hugo has received, amongst others, the DST-INSPIRE fellowship for his doctoral research and the Gold Medal from Pondicherry University for academic excellence during his Masters. His research has been published in high-impact peer-reviewed journals, including PLOS Neglected Tropical Diseases and Systematic Biology. When not working or writing, Hugo can be found consuming copious amounts of anime and manga, composing and making music with his bass guitar, shredding trails on his MTB, playing video games (he prefers the term ‘gaming’), or tinkering with all things tech.

Citations

Please use one of the following formats to cite this article in your essay, paper or report:

  • APA

    Francisco de Souza, Hugo. (2026, September 02). AI learned from cardiologists’ notes, and needed far less labeled ECG data. News-Medical. Retrieved on September 03, 2026 from https://www.news-medical.net/news/20260902/AI-learned-from-cardiologistse28099-notes-and-needed-far-less-labeled-ECG-data.aspx.

  • MLA

    Francisco de Souza, Hugo. "AI learned from cardiologists’ notes, and needed far less labeled ECG data". News-Medical. 03 September 2026. <https://www.news-medical.net/news/20260902/AI-learned-from-cardiologistse28099-notes-and-needed-far-less-labeled-ECG-data.aspx>.

  • Chicago

    Francisco de Souza, Hugo. "AI learned from cardiologists’ notes, and needed far less labeled ECG data". News-Medical. https://www.news-medical.net/news/20260902/AI-learned-from-cardiologistse28099-notes-and-needed-far-less-labeled-ECG-data.aspx. (accessed September 03, 2026).

  • Harvard

    Francisco de Souza, Hugo. 2026. AI learned from cardiologists’ notes, and needed far less labeled ECG data. News-Medical, viewed 03 September 2026, https://www.news-medical.net/news/20260902/AI-learned-from-cardiologistse28099-notes-and-needed-far-less-labeled-ECG-data.aspx.

Comments

The opinions expressed here are the views of the writer and do not necessarily reflect the views and opinions of News Medical.
Post a new comment
Post

While we only use edited and approved content for Azthena answers, it may on occasions provide incorrect responses. Please confirm any data provided with the related suppliers or authors. We do not provide medical advice, if you search for medical information you must always consult a medical professional before acting on any information provided.

Your questions, but not your email details will be shared with OpenAI and retained for 30 days in accordance with their privacy principles.

Please do not ask questions that use sensitive or confidential information.

Read the full Terms & Conditions.

You might also like...
Sunlight, blood sugar and diabetes: a new study finds an intriguing connection