spot_img
HomeResearch & DevelopmentDecoding AI's Medical Decisions: A New Dataset and LLM-Guided...

Decoding AI’s Medical Decisions: A New Dataset and LLM-Guided Explanations for Clinical Coding

TLDR: This research evaluates how well AI models explain their medical coding decisions, focusing on “faithfulness” (how accurately explanations reflect model reasoning) and “plausibility” (how well they match human expert judgment). The study introduces a new high-quality dataset for evaluating explanations and proposes using large language models (LLMs) to guide the learning of these explanations. They found that LLM-generated explanations align best with human experts, and providing a few human examples further improves these AI-generated explanations, though there’s a trade-off with the model’s overall coding accuracy.

Automated clinical coding, the process of translating free-text descriptions from Electronic Health Records (EHRs) into standardized codes like the International Classification of Diseases (ICD), plays a vital role in healthcare. These codes are crucial for billing, reimbursement, auditing, and supporting clinical decisions. Historically, this task relied on manual efforts, which were costly and prone to errors. Over time, rule-based systems, traditional machine learning, and more recently, deep learning models including advanced attention-based networks and transformers, have significantly improved efficiency and accuracy.

However, despite these advancements, a major challenge persists: the lack of explainability in these sophisticated AI models. Healthcare professionals and patients need to trust AI-driven outcomes, and this trust is undermined when model decisions are opaque. Researchers have been exploring ways to make these models more transparent, often by extracting or generating short text snippets, known as ‘rationales’, that highlight the evidence supporting a code assignment.

This research paper, titled “Evaluation and LLM-Guided Learning of ICD Coding Rationales” by Mingyang Li, Viktor Schlegel, Tingting Mu, Wuraola Oyewusi, Kai Kang, and Goran Nenadic, addresses critical gaps in the current understanding and generation of these explanations. The authors point out that existing evaluations of explainability often lack systematic analysis using consistent criteria on high-quality datasets, and there’s a need for dedicated approaches explicitly trained to generate better rationales. You can read the full paper here.

The study introduces a comprehensive evaluation framework for ICD coding rationales through two key perspectives: faithfulness and plausibility. Faithfulness assesses how accurately explanations reflect the model’s actual reasoning, while plausibility measures how consistent these explanations are with human expert judgment. To facilitate the evaluation of plausibility, the researchers constructed a new, high-quality rationale-annotated dataset called RD-IV-10. This dataset is built on the up-to-date MIMIC-IV benchmark with ICD-10 codes, offering denser annotations with diverse granularity and better alignment with current clinical practice, overcoming limitations found in previous datasets like MDACE.

Encouraged by the promising plausibility of rationales generated by Large Language Models (LLMs) for ICD coding, the paper further proposes novel rationale learning methods. These methods leverage rationales produced by prompting LLMs, with or without annotation examples, as distant supervision signals. The empirical findings are significant: LLM-generated rationales align most closely with those of human experts. Moreover, incorporating a small number of human-annotated examples (few-shot prompting) not only further improves the quality of LLM-generated rationales but also enhances rationale-learning approaches.

In their experiments, the authors compared the faithfulness of state-of-the-art ICD coding models (CAML, LAAT, and PLM-ICD) and found that PLM-ICD and LAAT generally performed better than CAML. For plausibility, LLM-generated rationales, particularly from Gemini 2-Flash, achieved the highest scores, outperforming both naive entity-level extractions and rationales derived from the models’ attention weights. While LLM-guided learning improved rationale plausibility, the study observed an interesting trade-off: a slight decrease in overall ICD coding performance when incorporating the additional learning objective for rationales. This suggests that balancing both high coding accuracy and strong rationale alignment remains a challenging task for future research.

Also Read:

This work highlights the critical importance of explainability in clinical AI and offers a robust framework and innovative methods for improving the quality of explanations in ICD coding. The new dataset and the insights into LLM-guided learning pave the way for more trustworthy and transparent AI applications in healthcare.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -