TLDR: A recent study by the Icahn School of Medicine at Mount Sinai has shown that generative artificial intelligence, specifically large language models like GPT-4, can significantly improve the accuracy of predicting patient admissions in emergency departments. The AI model, even with minimal training data, outperformed traditional machine learning models and offered transparent reasoning for its decisions, potentially streamlining clinical workflows and reducing variability in care.
Generative artificial intelligence (AI) is poised to revolutionize clinical decision-making within emergency departments (EDs), according to groundbreaking research from the Icahn School of Medicine at Mount Sinai. A study, published in the May 21 online issue of the Journal of the American Medical Informatics Association, demonstrates that large language models (LLMs) such as GPT-4 can effectively predict whether an emergency room patient requires hospital admission, even when trained on a limited number of records.
The motivation behind the research stemmed from the critical need to enhance the ability to predict admissions in high-volume settings like the ED. Dr. Eyal Klang, MD, co-senior author and Director of the Generative AI Research Program in the Division of Data-Driven and Digital Medicine (D3M) at Icahn Mount Sinai, stated, ‘Our goal is to enhance clinical decision-making through this technology. We were surprised by how well GPT-4 adapted to the ER setting and provided reasoning for its decisions. This capability of explaining its rationale sets it apart from traditional models and opens up new avenues for AI in medical decision-making.’
The retrospective study analyzed data from over 864,000 emergency room visits across seven Mount Sinai Health System hospitals. Researchers utilized both structured data, such as vital signs, and unstructured data, including nurse triage notes, while meticulously excluding identifiable patient information. Of these visits, 159,857, or 18.5 percent, resulted in hospital admissions.
GPT-4’s performance was rigorously compared against conventional machine-learning models, including Bio-Clinical-BERT for text analysis and XGBoost for structured data. The generative AI model was evaluated both independently and in combination with these traditional methods. A key finding was the LLMs’ ability to learn effectively from just a few examples, a stark contrast to traditional machine-learning models that often require millions of records for training. Furthermore, the study indicated that LLMs could incorporate predictions from traditional machine-learning methods, leading to improved overall performance.
Dr. Girish N. Nadkarni, MD, MPH, co-senior author, Irene and Dr. Arthur M. Fishberg Professor of Medicine at Icahn Mount Sinai, and System Chief of D3M, emphasized the broader implications: ‘This work opens the door for further innovation in health care AI, encouraging the development of models that can reason and learn from limited data, like human experts do.’ He added a crucial caveat, ‘However, while the results are encouraging, the technology is still in a supportive role, enhancing the decision-making process by providing additional insights, not taking over the human component of health care, which remains critical.’
Also Read:
- Landmark Study Reveals Generative AI Chatbots Currently Unreliable for Critical Stroke Care Advice
- Google’s Med-Gemini AI Identifies Non-Existent Brain Structure, Raising Medical Safety Concerns
The research suggests that AI could soon become a valuable support tool for doctors in emergency rooms, facilitating quick, informed decisions regarding patient admissions and potentially reducing variability in imaging exam orders by aligning with expert evidence-based guidelines. The team is now exploring further applications of LLMs in healthcare, aiming for seamless integration with existing methods to address complex clinical challenges in real-time.


