spot_img
HomeResearch & DevelopmentHealthBranches: A Dataset for Smarter Medical Question Answering

HealthBranches: A Dataset for Smarter Medical Question Answering

TLDR: HealthBranches is a new medical Q&A dataset designed to test complex reasoning in LLMs. It uses a semi-automated process to create 4,063 patient case studies across 17 healthcare topics, each with questions, answers, and explicit reasoning paths derived from medical decision pathways. This dataset helps evaluate LLMs’ multi-step inference and RAG performance, aiming for more trustworthy and clinically reliable AI in healthcare.

Large Language Models (LLMs) are transforming how we interact with information, including in healthcare. They can answer questions, summarize texts, and retrieve information, offering great potential for public health by providing insights on diseases and prevention. However, deploying these powerful AI models in medical settings comes with significant challenges. Issues like limited accuracy, inherent biases, data constraints, lack of contextual understanding, and the tendency to generate plausible but incorrect information (known as hallucinations) pose serious risks. Misleading medical information can directly impact patient health and even violate privacy regulations.

To address these critical limitations, researchers have explored strategies such as fine-tuning LLMs on domain-specific data, using Retrieval-Augmented Generation (RAG) to pull in external knowledge, and prompt engineering. RAG, in particular, has shown promise in reducing hallucinations and improving factual accuracy by integrating external domain knowledge. While RAG enhances LLM performance in medical question-answering, the precision of the retrieved content is paramount. Knowledge Graphs (KGs) can help by structuring domain-specific knowledge in a verifiable way, but there’s still a need for datasets that closely integrate structured medical knowledge with Q&A tasks, reflecting real-world diagnostic complexity.

Introducing HealthBranches: A New Medical Q&A Benchmark

A new research paper introduces HealthBranches, a novel benchmark dataset for medical Question-Answering (Q&A) that aims to evaluate complex reasoning in LLMs. This dataset is generated through a semi-automated pipeline that transforms explicit decision pathways from medical sources into realistic patient cases, complete with associated questions and answers. HealthBranches covers an impressive 4,063 case studies across 17 diverse healthcare topics, with each data point rooted in clinically validated reasoning chains.

What makes HealthBranches unique is its support for both open-ended and multiple-choice question formats, and crucially, it includes the full reasoning path for each Q&A pair. This structured design allows for a robust evaluation of LLMs’ multi-step inference capabilities, including their performance in structured RAG contexts. The goal is to lay a foundation for developing more trustworthy, interpretable, and clinically reliable LLMs for high-stakes medical domains, while also serving as a valuable educational resource.

How HealthBranches Was Built

The creation of HealthBranches involved a sophisticated semi-automated pipeline. It begins by parsing knowledge from medical textbooks, extracting both textual descriptions and graphical decision trees (knowledge graphs). These graphical representations formalize clinical decision-making processes. From these graphs, root-to-leaf traversals are enumerated, forming refined reasoning paths.

Next, an LLM (Gemini-flash 2.0) is prompted to generate Q&A content based on the extracted text and reasoning paths. It first produces an open-ended answer, which is then used as the correct option for a multiple-choice version, along with four plausible but incorrect distractors. To ensure quality and consistency, a two-stage refinement procedure is applied. Questions incorrectly answered by powerful models like Llama 3.1 are further examined and refined using GPT-4o with web search and reasoning capabilities, followed by a human reviewer audit. This hybrid approach, involving both Gemini and ChatGPT, helps reduce potential systematic biases from relying on a single model throughout the pipeline.

Key Features and Evaluation

Compared to existing medical Q&A benchmarks, HealthBranches stands out by uniquely combining multiple-choice and open-ended formats with explicit reasoning paths derived from medical knowledge graphs. While other datasets like MedQA and PubMedQA assess medical knowledge, they lack the explanation structure. MedCalc-BENCH provides explanations but focuses on numeric computation. HealthBranches, in contrast, enables the evaluation of non-computational, interpretable, and knowledge-driven reasoning, making it ideal for training and evaluating trustworthy medical LLMs.

The researchers evaluated 11 open, decoder-only LLMs, including Mistral, Llama, Gemma, and Qwen, under various settings: zero-shot (question only), RAG (with retrieved context), and “topline” settings (with only reasoning path, only textual description, or both). Evaluation metrics included Exact Match for quiz questions, an LLM-as-a-judge score (using Gemini-flash 2.0) for open-ended answers, and a semantic similarity score (using BGE-M3) to compare generated answers with ground truth.

Also Read:

Findings and Impact

The evaluation revealed that models significantly benefit from having the explicit reasoning path information, highlighting its value in guiding LLMs through complex medical decision-making. Interestingly, the improvement from RAG information was marginal for top-performing models, suggesting that larger, newer models might already incorporate some of this contextual knowledge. The consistency between quiz accuracy and open-ended answer quality (as assessed by the LLM-as-a-judge) further validates the dataset’s reliability.

HealthBranches represents a significant step forward in developing and assessing interpretable, safe, and trustworthy LLMs in medical settings. By aligning evaluations with clinical decision-making processes, it offers a valuable resource for both AI research and medical education. While the pipeline has limitations, such as potential biases from the LLMs used in its creation and the need for continued expert review, it provides a robust framework for advancing AI’s capabilities in healthcare. For more details, you can refer to the full research paper here.

Karthik Mehta
Karthik Mehtahttps://blogs.edgentiq.com
Karthik Mehta is a data journalist known for his data-rich, insightful coverage of AI news and developments. Armed with a degree in Data Science from IIT Bombay and years of newsroom experience, Karthik merges storytelling with metrics to surface deeper narratives in AI-related events. His writing cuts through hype, revealing the real-world impact of Generative AI on industries, policy, and society. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -