TLDR: This research introduces a novel computational method using Latent Dirichlet Allocation (LDA) and Large Language Models (LLMs) to analyze and summarize lived healthcare experiences from African American individuals’ stories. By identifying 26 key topics such as chronic pain management and caregiving, the approach provides detailed insights into healthcare disparities and potential intervention areas. The study demonstrates that AI can efficiently process unstructured narrative data, with LLM-generated summaries showing high accuracy, comprehensiveness, and usefulness, validated against human expert assessments. This work offers a promising avenue for improving health outcomes and equity by leveraging the communicative power of storytelling.
Storytelling is a fundamental way humans communicate, offering deep insights into personal beliefs and emotions that go beyond simple facts. In healthcare, these narratives can reveal crucial factors contributing to disparities in health outcomes and suggest new ways to intervene and improve care. However, analyzing vast amounts of unstructured spoken stories, like those from patient experiences, has traditionally been a time-consuming and expensive task.
A recent research paper, Topic-aware Large Language Models for Summarizing the Lived Healthcare Experiences Described in Health Stories, introduces an innovative computational method that leverages artificial intelligence to efficiently analyze these rich narratives. The study focuses specifically on the healthcare experiences of African American storytellers, a population that often faces worse health outcomes and reduced access to services compared to white individuals in the United States.
Bridging the Gap with AI
The researchers aimed to determine if Large Language Models (LLMs) could identify underlying factors and potential avenues for intervention by performing topic-aware hierarchical summarization of these health stories. They utilized a dataset of fifty transcribed stories from African American individuals, drawn from the MyPaTH Story Booth archive, a collection of over 1,500 personal healthcare narratives.
The methodology involved a multi-step process. First, the Latent Dirichlet Allocation (LDA) technique, a classical topic modeling approach, was used to identify recurring topics within the stories. This step helped to categorize the diverse experiences shared by participants. To make these topics clinically interpretable, an open-source LLM (LLaMA-3.1) was then employed to generate clear labels for each topic.
Following topic identification and labeling, the core of the research involved a hierarchical summarization approach powered by LLMs. This technique first generated individual summaries for each story related to a specific topic. Then, these individual story summaries were further summarized to create a comprehensive topic summary, effectively condensing long-form narratives into digestible insights while overcoming the context length limitations often faced by LLMs.
Key Findings and Insights
The study successfully identified 26 distinct topics from the fifty African American stories. These topics covered a wide range of experiences, including ‘health behaviors,’ ‘interactions with medical team members,’ ‘caregiving and symptom management,’ ‘chronic pain management,’ ‘doctor-patient relationship,’ and ‘hospital experience,’ among others. The LLM-generated topic summaries were evaluated for fabrication, accuracy, comprehensiveness, and usefulness using GPT-4 Turbo as an evaluator, a method validated against human expert assessments.
The evaluation showed that the topic summaries were largely free from fabrication, highly accurate, comprehensive, and useful, with moderate to high agreement between GPT-4 ratings and expert assessments. This suggests that AI can reliably process and summarize complex qualitative data, offering a powerful tool for researchers and healthcare professionals.
For instance, the summary for ‘Chronic Pain Management’ highlighted participants’ struggles with inadequate pain management, ineffective treatments, and feelings of not being taken seriously by providers. It also noted the use of alternative methods like CBD oil and yoga, and the benefits of supportive doctors and pain clinics. Similarly, the ‘Caregiving Experience’ summary revealed challenges such as lack of support, financial struggles, and emotional toll, while emphasizing the importance of compassion and the need for respite care and financial assistance.
Also Read:
- Generating Realistic Aphasia Transcripts with AI: A New Approach to Data Scarcity
- AI Models Chart Unexplored Territories in Biomedical Research by Identifying Knowledge Gaps
Implications and Future Directions
This approach offers a significant advancement in understanding healthcare experiences beyond traditional sentiment analysis, providing fine-grained details about clinical interactions, treatment efficacy, and their broader impact on daily life. The ability to efficiently extract such insights from large narrative datasets can help identify potential factors contributing to health disparities and inform targeted interventions to improve the healthcare ecosystem.
While the study demonstrated promising results, the authors acknowledge limitations such as potential transcription errors and some redundancy among identified topics. Future work aims to explore larger datasets, compare different LLMs for summarization, and investigate end-to-end AI approaches that integrate topic modeling and summarization more seamlessly. Ultimately, this research paves the way for leveraging computational methods to enhance health equity and improve patient and caregiver support by truly listening to and learning from lived experiences.


