spot_img
HomeResearch & DevelopmentBoosting Trust in AI-Generated Liver MRI Reports: A New...

Boosting Trust in AI-Generated Liver MRI Reports: A New Framework for Prompt Optimization and Credibility Assessment

TLDR: A study introduces a Multi-Dimensional Credibility Assessment (MDCA) framework to evaluate the trustworthiness of LLM-generated liver MRI reports across semantic coherence, diagnostic correctness, and clinical prioritization. It also provides guidance on prompt optimization, finding that a combination of structured instructions and 10-15 examples significantly improves report quality. Kimi-K2 and DeepSeek-V3 were identified as top-performing models, demonstrating that thoughtful prompt engineering is key to integrating LLMs safely and effectively into radiology.

Large language models (LLMs) are showing great promise in helping radiologists create diagnostic reports from imaging findings. These AI tools could assist with radiology reporting, training new professionals, and ensuring quality control. However, there hasn’t been much clear guidance on how to best design prompts for these LLMs in different clinical situations, nor a comprehensive way to assess how trustworthy their generated reports are.

A recent study by Qiuli Wang, Jie Cheng, Yongxu Liu, Xingpeng Zhang, Xiaoming Li, and Wei Chen addresses these challenges. Their research, titled “From Prompt Optimization to Multi-Dimensional Credibility Evaluation: Enhancing Trustworthiness of Chinese LLM-Generated Liver MRI Reports,” introduces a new framework to improve the reliability of LLM-generated liver MRI reports. You can read the full paper here: Research Paper.

A New Framework for Trustworthiness

The study proposes a Multi-Dimensional Credibility Assessment (MDCA) framework. This framework is designed to evaluate the trustworthiness of LLM-generated radiology reports across three key areas:

  • Semantic Coherence (SC): This checks if the report reads smoothly and follows typical radiological writing styles. It focuses on the linguistic flow and consistency.
  • Diagnostic Correctness (DC): This measures the accuracy and completeness of the diagnostic information in the report, scoring based on disease coverage and precise terminology.
  • Clinical Prioritization Alignment (CPA): This assesses whether the most urgent or important findings are presented at the beginning of the report, mirroring real-world clinical priorities. For instance, malignant tumors should appear first.

The final MDCA score combines these three dimensions, with Diagnostic Correctness and Clinical Prioritization Alignment given higher importance (40% each) compared to Semantic Coherence (20%), as generating fluent text is generally easier for LLMs than ensuring diagnostic accuracy.

Optimizing Prompts for Better Performance

The researchers also provided guidance on how to optimize prompts for LLMs in specific institutional contexts. They compared eleven different prompt configurations, which included various combinations of instructional components and example-based guidance. Instruction-based components covered aspects like defining the LLM’s role (e.g., a radiologist with 30 years of experience), specifying the core task, using a tiered diagnostic taxonomy (TOP system), including mandatory verification items, setting report structure standards, and incorporating imaging diagnostic principles. Example-based guidance involved providing sample reports written by experienced radiologists.

The study used several advanced Chinese LLMs, including Kimi-K2-Instruct-0905 (Moonshot AI), Qwen3-235B-A22B-Instruct-2507 (Alibaba Group), DeepSeek-V3 (DeepSeek AI), and ByteDance-Seed-OSS-36B-Instruct (ByteDance), all evaluated on the SiliconFlow platform using over 15,000 institutional liver cancer reports.

Key Findings

The MDCA framework proved effective in evaluating the quality of LLM-generated reports. The research revealed that optimizing prompt design significantly improved diagnostic accuracy and clinical prioritization. Specifically:

  • Prompts with role definition but lacking explicit diagnostic guidance performed poorly.
  • Including example-based guidance substantially improved overall performance, especially semantic coherence, making outputs more fluent and logically consistent.
  • The optimal balance between performance and efficiency was achieved with approximately ten examples, particularly when computational resources were limited.
  • Integrating key instruction-based components (like the tiered diagnostic taxonomy and mandatory verification items) further enhanced semantic coherence, diagnostic correctness, and top-1 matching scores, leading to more accurate reports.

Among the LLMs tested, Kimi-K2-Instruct-0905 and DeepSeek-V3 consistently delivered the best overall performance. Kimi-K2 achieved the highest composite score of 76.149, while DeepSeek-V3 followed closely with 75.410, both demonstrating stable and reliable results. The study also found that the effect of example numbers was consistent across models, suggesting that these performance gains come from improved contextual understanding rather than model-specific traits.

Also Read:

Implications for Radiology

This research highlights that institution-specific prompt optimization and multi-dimensional credibility evaluation are crucial for enhancing the trustworthiness of LLM-generated liver MRI reports. The proposed frameworks not only boost reporting reliability and interpretability but also offer practical tools for quality control in radiology and for training new radiologists. This supports the safe and standardized integration of LLMs into clinical workflows, ultimately benefiting patient care by improving the accuracy and consistency of diagnostic reports.

Nikhil Patel
Nikhil Patelhttps://blogs.edgentiq.com
Nikhil Patel is a tech analyst and AI news reporter who brings a practitioner's perspective to every article. With prior experience working at an AI startup, he decodes the business mechanics behind product innovations, funding trends, and partnerships in the GenAI space. Nikhil's insights are sharp, forward-looking, and trusted by insiders and newcomers alike. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -