spot_img
HomeResearch & DevelopmentAdvancing Heart Health Prediction with AI: Integrating Genetics and...

Advancing Heart Health Prediction with AI: Integrating Genetics and ECG Data Using Language Models

TLDR: This research introduces a new AI framework that uses large language models to predict cardiovascular disease risk. It combines genetic information (SNP variants) and heart activity data (ECG phenotypes) even when very few patient labels are available. The system uses a clever “pseudo-labeling” method to learn from limited data and provides “Chain of Thought” explanations to make its predictions understandable, showing improved accuracy and generalizability compared to traditional methods.

Cardiovascular disease (CVD) remains a leading cause of death globally, making early and accurate risk assessment incredibly important. However, predicting who is at risk is a complex challenge because many factors contribute to CVD, and high-quality, labeled patient datasets are often scarce. While genetic information, like SNP variants, and heart activity data, such as ECG phenotypes, are becoming more accessible, effectively combining these different types of data, especially when there are few labeled examples, has been difficult.

This research introduces a groundbreaking approach to tackle this problem: a few-label multimodal framework that uses large language models (LLMs) to integrate genetic and electrophysiological information for cardiovascular risk stratification. The core idea is to leverage the powerful pattern-recognition capabilities of LLMs to make sense of diverse biological signals, even with limited direct supervision.

A Novel Approach to Data Scarcity and Interpretability

The framework addresses the scarcity of well-annotated multimodal datasets by incorporating a ‘pseudo-label refinement’ strategy. This method adaptively identifies high-confidence labels from weakly supervised predictions, allowing the model to be robustly fine-tuned with only a small set of actual ground-truth annotations. This is crucial in biomedical fields where obtaining extensive, curated labels is often costly and time-consuming.

To enhance understanding and trust in the model’s predictions, the task is framed as a ‘Chain of Thought’ (CoT) reasoning problem. This means the model is prompted to produce clinically relevant explanations alongside its risk predictions. This capability is vital for clinical decision-making, as it provides transparent insights into how genetic and physiological features contribute to a patient’s specific cardiovascular risk.

How the System Works

The researchers constructed a harmonized cardiogenomic dataset by integrating high-resolution SNP genotyping with morphological and temporal ECG features. This dataset was curated from the PhenoAI HPP repository. Participants were categorized into three tiers based on the availability of clinical labels:

  • Tier 1: Individuals with confirmed cardiac diagnoses.
  • Tier 2: Participants with indirect cardiovascular risk factors (e.g., hypertension).
  • Tier 3: Unlabeled participants with no known prior cardiac diagnosis.

The process involves several stages:

First, ‘pseudo-labels’ are generated across all three tiers. SNP and ECG features are preprocessed and then embedded into dense vector representations. A clustering algorithm (k-means) is applied to these combined representations to derive latent genotype–phenotype groupings, assigning a pseudo-label to each participant.

Next, a pretrained transformer model (a domain-adapted BioBERT model) is fine-tuned on a subset of these pseudo-labels. The most reliable clusters are then selected based on predictive consistency, forming a refined pseudo-labeled dataset. This refined data is used to further fine-tune the language model.

Finally, the model is adapted to a small set of gold-standard clinical labels using parameter-efficient fine-tuning (LoRA). During this phase, Chain-of-Thought prompting guides the model’s reasoning, integrating ECG features, disease-associated SNPs, and tier-specific risk annotations to infer potential cardiovascular risk.

Key Findings and Performance

The study evaluated three base causal LLMs: GPT-2, DeepSeek 1.3B, and Llama 3.2 1B. Experimental results demonstrated that integrating multimodal inputs, few-label supervision, and CoT reasoning significantly improves robustness and generalizability across diverse patient profiles.

An ablation study, where either SNP or ECG inputs were removed, showed a significant performance drop across all metrics. This confirms that the fusion of genomic and ECG features works synergistically for cardiovascular risk prediction. ECG inputs were particularly important for identifying subtle risk indicators (recall), while SNP information enhanced the model’s ability to differentiate true positives (precision).

DeepSeek 1.3B consistently achieved the highest overall performance, with an accuracy of 0.910, precision of 0.869, recall of 0.830, and an F1 score of 0.840. This suggests that DeepSeek is highly effective in classifying samples and maintaining a balanced trade-off between precision and recall. DeepSeek also exhibited the most rapid convergence during training, indicating efficient adaptation to the task data.

The research also found that the performance gap between models trained with full supervision (Skyline) and those trained with few labels was notably small for DeepSeek 1.3B, often less than 0.02 in F1 score. This highlights its strong generalization capabilities even with limited supervision. As the number of available labels increased from 50 to 350, all models showed improved performance, with DeepSeek and Llama demonstrating rapid gains, underscoring the importance of label quantity in stabilizing training for larger models.

Also Read:

Implications for Personalized Cardiovascular Care

This study underscores the promise of LLM-based few-label multimodal modeling for advancing personalized cardiovascular care. By effectively integrating genetic and electrophysiological data, the framework offers a deeper understanding of the molecular and electrophysiological mechanisms underlying cardiovascular disease. The resulting multimodal embeddings may reflect an emergent representation of cardiac physiology, where genomic variants shape electrophysiological expression patterns observable through ECG signals.

While the current methodology focuses on a subset of SNP and ECG features and acknowledges potential noise propagation in pseudo-labeling, the proposed few-labels approach represents a viable route towards practical deployment in low-resource clinical settings. Future work could incorporate additional data sources like imaging, clinical notes, or biochemical data, and scale to larger LLM architectures.

Overall, this framework provides a scalable foundation for next-generation clinical decision-support systems that balance predictive performance with transparency, a critical requirement for AI deployment in medicine. For more details, you can read the full research paper here.

Karthik Mehta
Karthik Mehtahttps://blogs.edgentiq.com
Karthik Mehta is a data journalist known for his data-rich, insightful coverage of AI news and developments. Armed with a degree in Data Science from IIT Bombay and years of newsroom experience, Karthik merges storytelling with metrics to surface deeper narratives in AI-related events. His writing cuts through hype, revealing the real-world impact of Generative AI on industries, policy, and society. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -