spot_img
HomeResearch & DevelopmentUnpacking Multimodal Learning in Spatial Transcriptomics with HESCAPE

Unpacking Multimodal Learning in Spatial Transcriptomics with HESCAPE

TLDR: HESCAPE is a new large-scale benchmark for evaluating how well AI models can combine histology images and gene expression data in spatial transcriptomics. It found that gene expression encoders are key for alignment, and while cross-modal pretraining helps classify gene mutations, it surprisingly hinders direct gene expression prediction, largely due to batch effects in the data. HESCAPE is open-sourced to encourage further research into batch-robust multimodal learning.

Spatial transcriptomics is a groundbreaking field that allows scientists to simultaneously measure gene activity and observe tissue structure. This provides an unparalleled view into how cells are organized and how diseases develop. However, a significant challenge in this area has been the lack of comprehensive benchmarks to evaluate how well different computational methods can learn from both histology images and gene expression data together.

Introducing HESCAPE: A New Benchmark

To address this gap, researchers have introduced HESCAPE, a large-scale benchmark designed for cross-modal contrastive pretraining in spatial transcriptomics. HESCAPE is built upon a carefully curated dataset that spans multiple organs, includes 6 different gene panels, and covers 54 donors. This extensive dataset provides a robust foundation for evaluating multimodal learning approaches.

The benchmark systematically assessed various state-of-the-art image and gene expression encoders using multiple pretraining strategies. Their effectiveness was then tested on two crucial downstream tasks: classifying gene mutations and predicting gene expression.

Key Findings from the Benchmark

One of the primary findings from HESCAPE is that gene expression encoders play the most critical role in achieving strong representational alignment between the two data types. Specifically, gene models that were pretrained on spatial transcriptomics data consistently outperformed those trained without spatial data, as well as simpler baseline methods. The DRVI gene encoder, when paired with image encoders like Gigapath, H0mini, and UNI, showed the best performance in cross-modal retrieval tasks (matching images to genes and vice versa).

Interestingly, while large-scale gene foundation models exist, they did not surpass VAE-based models like DRVI when the latter were pretrained on the specific dataset of interest. However, these foundation models still showed substantial improvements over basic MLP baselines.

Downstream Task Performance: A Nuanced Picture

The evaluation of downstream tasks revealed a more complex outcome. For gene mutation classification, contrastive pretraining consistently improved performance. For example, in colorectal cancer, HESCAPE-trained models showed significant gains in predicting MSI and BRAF mutations. Similar improvements were observed for ER and PR status in breast cancer, and EGFR and KRAS prediction in lung cancer. This suggests that integrating molecular knowledge into vision models can indeed enhance the detection of certain biomarkers from tissue images.

However, a striking contradiction emerged when it came to direct gene expression prediction from histology images. Surprisingly, contrastive pretraining often degraded performance compared to baseline encoders that were trained without cross-modal objectives. This counterintuitive result suggests that the alignment process might force the image encoder to discard valuable morphological and spatial information essential for accurate gene expression prediction.

The Role of Batch Effects

The researchers identified batch effects as a key factor interfering with effective cross-modal alignment. Batch effects are technical variations introduced during data collection or processing, which can obscure true biological signals. The study found a clear relationship: datasets with better batch integration (meaning fewer technical variations) consistently achieved superior cross-modal retrieval performance. This highlights the critical need for developing batch-robust multimodal learning approaches in spatial transcriptomics.

Also Read:

Conclusion and Future Directions

HESCAPE provides a valuable resource for the scientific community, offering standardized datasets, evaluation protocols, and benchmarking tools. The findings underscore the importance of pretraining gene encoders on relevant spatial transcriptomics data and highlight the complex interplay between cross-modal alignment and downstream task performance. The unexpected degradation in gene expression prediction suggests that future research should focus on developing multimodal encoders that can explicitly account for and mitigate domain-specific effects, such as batch effects, to learn more robust and generalizable representations. This work is a significant step towards advancing both fundamental multi-omics research and translational applications in digital pathology and patient stratification. You can find more details about HESCAPE and access the resources at the project’s GitHub repository.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -