spot_img
HomeResearch & DevelopmentText-Guided AI Improves Lesion Detection in CT Scans

Text-Guided AI Improves Lesion Detection in CT Scans

TLDR: A new AI model, Text-Swin-UMamba, integrates short text descriptions from radiology reports with CT imaging data to improve lesion segmentation. Tested on the ULS23 DeepLesion dataset, it significantly outperformed previous text-embedding models and showed better or comparable performance to purely image-based models, demonstrating the benefit of combining textual and visual information for more accurate clinical assessments.

In the evolving landscape of medical diagnostics, the accurate and efficient assessment of lesions on CT scans is crucial for managing chronic diseases like cancer and lymphoma. Traditionally, radiologists manually measure the size and characteristics of lesions, a process that can be time-consuming and challenging, especially with numerous lesions or variations in scanner types and lesion appearances.

A new research paper introduces an innovative approach to automate and enhance this process by integrating textual descriptions from radiology reports directly into the image segmentation workflow. The study, titled “Text Embedded Swin-UMamba for DeepLesion Segmentation,” explores the feasibility of combining imaging features with descriptive text to improve the accuracy of lesion segmentation.

The core of this advancement is the proposed Text-Swin-UMamba model. Unlike previous methods that primarily rely on image data or struggle to adapt large language models (LLMs) to the nuances of medical imaging, Text-Swin-UMamba is designed to seamlessly blend visual and textual information. It builds upon the robust Swin-UMamba architecture, which is known for its effectiveness in medical image segmentation, by adding a ‘Text Tower’ encoder.

The Text Tower is a specialized component that processes short-form text descriptions from radiology reports. It tokenizes the text, passes it through a pre-trained ‘BioLord’ model (which is adept at understanding clinical sentences), and then generates a summarized, fixed-size embedding representing the semantic meaning of the text. This textual embedding is then intelligently integrated into multiple stages of the Swin-UMamba image decoder. This multi-scale fusion allows the model to leverage the clinical context provided by the text, guiding the image segmentation process for better results.

The researchers evaluated Text-Swin-UMamba using the publicly available ULS23 DeepLesion dataset, which includes CT image slices and corresponding short-form descriptions of findings. The dataset is particularly challenging due to the small or tiny regions occupied by many lesion masks within the 512×512 pixel images.

The results were highly promising. On the test dataset, Text-Swin-UMamba achieved a high Dice Score of 82% and a low Hausdorff distance of 6.58 pixels, indicating excellent segmentation accuracy. Notably, the model demonstrated a significant 37% improvement in Dice score over a prior LLM-driven model, LanGuideMedSeg. It also surpassed purely image-based models like xLSTM-UNet and nnUNet, showing consistent gains in performance.

This study highlights the substantial advantage of incorporating short-form text descriptions into lesion segmentation. The clinical context derived from both imaging and text features helps the network achieve superior segmentation results. While the current work focuses on 2D images and a specific dataset design, the researchers note that the text embedding mechanism can be applied to other nnU-Net models and their derivatives. Future work aims to extend this approach to full-size 3D CT volumes and enable the use of long-form descriptions, potentially enhancing the explainability of lesion segmentation.

Also Read:

The dataset and code for this innovative model have been open-sourced, encouraging further research and development in this critical area of medical AI. You can find more details about the research paper here.

Ananya Rao
Ananya Raohttps://blogs.edgentiq.com
Ananya Rao is a tech journalist with a passion for dissecting the fast-moving world of Generative AI. With a background in computer science and a sharp editorial eye, she connects the dots between policy, innovation, and business. Ananya excels in real-time reporting and specializes in uncovering how startups and enterprises in India are navigating the GenAI boom. She brings urgency and clarity to every breaking news piece she writes. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -