TLDR: A new study evaluates GPT-5’s performance in medical reasoning, showing it consistently outperforms GPT-4o and other smaller models on text-based and multimodal medical question-answering benchmarks like MedQA, USMLE, and MedXpertQA. Notably, GPT-5 surpasses pre-licensed human experts in multimodal reasoning, demonstrating significant advancements in integrating visual and textual medical information for clinical decision support.
Recent advancements in large language models (LLMs) are transforming various fields, and medicine is no exception. Medical decision-making often requires integrating diverse information, from patient notes and structured data to complex medical images. A new study explores the capabilities of GPT-5, positioning it as a versatile multimodal reasoner for medical decision support.
The research, titled “Capabilities of GPT-5 on Multimodal Medical Reasoning,” was conducted by Shansong Wang, Mingzhe Hu, Qiang Li, Mojtaba Safari, and Xiaofeng Yang from the Department of Radiation Oncology, Winship Cancer Institute, Emory University School of Medicine. Their work systematically evaluates GPT-5’s performance in zero-shot chain-of-thought reasoning across both text-based question answering and visual question answering tasks.
Evaluating GPT-5’s Medical Prowess
To assess GPT-5, the researchers benchmarked it against several established datasets, including MedQA, MedXpertQA (both text and multimodal versions), MMLU medical subsets, USMLE self-assessment exams, and VQA-RAD. They compared GPT-5 with GPT-5-mini, GPT-5-nano, and GPT-4o-2024-11-20, using a unified protocol to ensure fair comparisons. This involved a “zero-shot chain-of-thought” approach, where the model was prompted to think step-by-step before providing a final answer, without prior fine-tuning for specific tasks.
Impressive Gains in Text and Multimodal Reasoning
The results were compelling. On text-based medical question answering, GPT-5 consistently outperformed all baselines. For instance, on MedQA (US 4-option), it achieved 95.84% accuracy, a notable 4.80% improvement over GPT-4o. The most significant gains were observed in MedXpertQA Text, where GPT-5 improved reasoning accuracy by 26.33% and understanding by 25.30% compared to GPT-4o. This suggests a substantial enhancement in the model’s ability to perform multi-step inference and comprehend complex medical narratives.
GPT-5 also demonstrated superior performance on the USMLE Self Assessment exams, surpassing all baselines across all three steps. Its average score across these steps reached 95.22%, exceeding typical human passing thresholds by a significant margin, indicating its readiness for high-stakes clinical reasoning tasks.
In the realm of multimodal reasoning, which involves integrating visual and textual information, GPT-5 showed a dramatic leap. On MedXpertQA MM, it achieved reasoning and understanding gains of +29.62% and +36.18% respectively, relative to GPT-4o. This highlights a significantly enhanced capability in combining visual and textual cues for diagnosis and decision-making. While GPT-5 scored slightly lower on VQA-RAD compared to GPT-5-mini, this was attributed to the dataset’s smaller scale and specific nature, possibly leading to more conservative reasoning in the larger model.
Surpassing Human Experts
Perhaps the most striking finding is GPT-5’s performance compared to pre-licensed human experts. While GPT-4o often performed below human experts, GPT-5 not only closed this gap but substantially surpassed them. For example, in MedXpertQA MM, GPT-5 improved multimodal reasoning by +24.23% and understanding by +29.40% over human experts. This suggests that GPT-5’s unified vision-language reasoning pipeline can integrate textual and visual evidence in a way that even experienced clinicians might find challenging under time-limited test conditions.
Also Read:
- MultiMedEdit: A New Benchmark for Updating Medical AI Knowledge
- Orchestrating Open-Source AI for Enhanced Medical Diagnosis
Future Implications for Medical AI
These findings mark a significant advancement in the capabilities of large language models for medical applications. GPT-5’s ability to consistently outperform baselines and even human experts on standardized benchmarks suggests its strong potential as a core component for future clinical decision-support systems. It can integrate complex textual and visual information streams to produce accurate, well-justified recommendations.
However, the researchers caution that these evaluations were conducted under idealized testing conditions and do not fully capture the complexity, uncertainty, and ethical considerations of real-world medical practice. Future work will need to explore prospective clinical trials and adaptive strategies to ensure safe and transparent deployment of such powerful AI tools. For more details, you can read the full research paper here.


