TLDR: A new method combines deep learning with expert human review to accurately segment vestibular schwannoma tumors on MRI scans. This “human-in-the-loop” approach iteratively refines an AI model, significantly improving segmentation accuracy and efficiency by 37.4% compared to traditional manual annotation. The resulting dataset, featuring 190 patients and 534 annotated scans, is publicly available to advance medical imaging research.
Accurate identification and measurement of vestibular schwannoma (VS) tumors on Magnetic Resonance Imaging (MRI) scans are crucial for patient care. Traditionally, this process relies on time-consuming manual annotations by highly skilled experts. While deep learning (DL) has shown promise in automating this task, achieving consistent and robust performance across varied datasets and complex clinical scenarios remains a significant challenge.
A recent research paper introduces an innovative approach that combines the power of deep learning with the invaluable insights of human experts to create a high-quality, annotated dataset for VS segmentation. This “human-in-the-loop” framework aims to make the annotation process more efficient and reliable, ultimately leading to better diagnostic tools and patient management.
The Challenge of Medical Image Annotation
Annotating medical data, especially for complex structures like VS tumors, is inherently difficult. Tumors vary greatly in size and shape, and anatomical differences between patients add to the complexity. For VS, precise volumetric measurement is key to tracking tumor growth, but even experts can have variability in their annotations, particularly for smaller tumors. Existing automated models often struggle to generalize across the diverse MRI hardware and imaging protocols used in routine clinical practice, necessitating new methods for creating trustworthy and representative datasets.
A Collaborative Approach: AI Meets Expert Review
The researchers developed a bootstrapped deep learning framework that iteratively refines VS segmentation. This framework integrates three main components:
- A 3D nnUNet model for automated VS segmentation, which improves its accuracy with each round of training.
- A multi-round quality assessment process involving expert review and manual corrections, driven by consensus among multiple specialists.
- An expert-driven validation process on diverse datasets to ensure the model’s generalizability and robustness.
The process begins with an initial deep learning model generating segmentation proposals. These proposals are then reviewed by a team of independent experts, including neurosurgical fellows and a trained radiologist. Annotations are categorized as ‘Accept,’ ‘Reject,’ or ‘Other’ (requiring further discussion). Complex cases flagged as ‘Other’ proceed to a consensus meeting involving the independent experts and a consultant neurosurgeon, where decisions are finalized.
Accepted annotations are used to retrain and refine the deep learning model in a process called bootstrapping. Rejected cases are re-processed by the improved model, and new proposals are reviewed and corrected by a trained radiologist. This iterative cycle of AI-generated proposals, expert review, and model refinement ensures that the dataset becomes increasingly accurate and reliable.
Significant Improvements in Accuracy and Efficiency
The proposed framework demonstrated a significant improvement in segmentation accuracy. The Dice Similarity Coefficient (DSC), a common metric for evaluating segmentation performance, increased from 0.9125 to 0.9670 on the internal validation dataset. This indicates a substantial enhancement in how well the automated segmentations match expert annotations. Importantly, the model maintained stable performance on external datasets, suggesting its ability to generalize to new clinical settings.
Beyond accuracy, the approach also delivered impressive efficiency gains. The total time required for the human-in-the-loop annotation process for 427 cases was approximately 6.35 hours. In contrast, traditional manual annotation of the same dataset would have taken about 10.15 hours, representing a remarkable 37.4% reduction in time. This efficiency is particularly evident in the time saved during quality assessment and correction, especially for larger tumors.
Also Read:
- Uncovering Age Bias in Medical Image Segmentation: A Deep Dive into Label and Representational Disparities
- Benchmarking Privacy and Performance in Federated Tumor Segmentation
A Valuable Public Resource
The finalized dataset includes 190 patients, with tumor annotations available for 534 longitudinal contrast-enhanced T1-weighted (T1CE) scans from 184 patients. This comprehensive dataset also includes demographic and clinical information, providing a rich resource for further research. The dataset is publicly accessible on The Cancer Imaging Archive (TCIA) at https://doi.org/10.7937/bq0z-xa62, fostering open science and accelerating advancements in medical imaging.
This work highlights the potential of integrating AI with expert human knowledge to create robust and efficient solutions for complex medical tasks. By combining automated segmentation with rigorous expert validation, the researchers have developed a clinically adaptable and generalizable strategy for automated VS segmentation, paving the way for improved patient management in diverse clinical environments.


