TLDR: A study evaluated GPT-5 and GPT-4o on mammogram visual question answering (VQA) across four public datasets. GPT-5 consistently outperformed other GPT variants but significantly lagged behind human experts and specialized AI models in accuracy, sensitivity, and specificity for tasks like BI-RADS assessment and malignancy classification. While not ready for clinical use, the performance improvement from GPT-4o to GPT-5 shows promise for general large language models in mammography VQA.
Mammogram visual question answering (VQA) is an innovative approach that combines image interpretation with clinical reasoning, holding significant promise for enhancing breast cancer screening. Interpreting mammograms is a complex and time-consuming task, even for highly experienced radiologists, due to the subtle nature of early-stage cancers and the variability in breast tissue patterns. This inherent difficulty has driven the development of artificial intelligence (AI) technologies to support radiologists, aiming to improve diagnostic accuracy and consistency.
A recent research paper titled “Is ChatGPT-5 Ready for Mammogram VQA?” explores the capabilities of the GPT-5 family and the GPT-4o model in this critical medical domain. The study, conducted by Qiang Li, Shansong Wang, Mingzhe Hu, Mojtaba Safari, Zachary Eidex, and Xiaofeng Yang from the Department of Radiation Oncology, Winship Cancer Institute, Emory University School of Medicine, systematically evaluated these models on four widely recognized public mammography datasets: EMBED, InBreast, CMMD, and CBIS-DDSM. The evaluation focused on key tasks such as BI-RADS assessment, abnormality detection, and malignancy classification.
Methodology and Evaluation
To assess GPT-5’s performance, the researchers created clinically relevant VQA items from structured labels within the datasets. These questions targeted specific clinical tasks, such as identifying BI-RADS breast density or determining if an imaging finding suggests malignancy. The study employed a zero-shot chain-of-thought (CoT) prompting strategy, meaning the models were tested without any specific fine-tuning on mammography data. This approach aimed to evaluate the models’ inherent, out-of-the-box capabilities in interpreting mammograms and answering clinically framed questions.
Key Findings
The evaluation revealed that GPT-5 consistently outperformed its smaller variants (GPT-5-mini, GPT-5-nano) and the previous generation GPT-4o across all datasets. For instance, on the EMBED dataset, GPT-5 achieved the highest scores among GPT models in density (56.8%), distortion (52.5%), mass (64.5%), calcification (63.5%), and malignancy (52.8%) classification. On InBreast, it attained 36.9% BI-RADS accuracy, 45.9% abnormality detection, and 35.0% malignancy classification. For CMMD, GPT-5 reached 32.3% abnormality detection and 55.0% malignancy accuracy. On CBIS-DDSM, it achieved 69.3% BI-RADS accuracy, 66.0% abnormality detection, and 58.2% malignancy accuracy.
However, despite these promising results, GPT-5’s performance still lagged significantly behind both human experts and specialized, fine-tuned AI models designed specifically for mammography. For example, when compared to human expert estimations on the CBIS-DDSM dataset, GPT-5 exhibited lower sensitivity (63.5% vs. 86.9%) and specificity (52.3% vs. 88.9%). This gap highlights the challenges general-purpose AI systems face in tasks requiring highly nuanced visual discrimination and integration of complex clinical context.
Understanding Model Limitations
Case studies of GPT-5’s errors provided further insights. The model sometimes misclassified extremely dense breasts (BI-RADS D) as heterogeneously dense (BI-RADS C), indicating a tendency to underestimate overall density in very dense cases. It also showed a propensity for false-positive interpretations, misclassifying benign structural changes as malignant, especially when confronted with architectural distortion or irregular mass margins without other clear malignant features. This suggests that without targeted domain adaptation and optimization, current general-purpose multimodal large language models (LLMs) are not yet sufficient for high-stakes clinical imaging applications.
Also Read:
- OpenAI’s O3 Model Surpasses Newer GPT-5 in Complex Multi-Application Office Workflows
- Beyond Solving: How Large AI Models Learn to Seek Information
Future Outlook
While GPT-5 is not yet ready for direct clinical use in mammography VQA, the study concludes with an optimistic outlook. The substantial performance improvements observed from GPT-4o to GPT-5 demonstrate a promising trend in the potential for general LLMs to assist with mammography VQA tasks. Further research focusing on domain-specific fine-tuning, reasoning transparency, and uncertainty calibration will be crucial for these models to achieve expert-level accuracy and reliability in medical imaging. For more detailed information, you can refer to the full research paper available at arXiv.org.


