TLDR: TalkDep is a novel system that uses advanced large language models (LLMs) to create clinically validated, realistic virtual patients for training and evaluating depression screening tools. It addresses the shortage of real patient data by generating diverse patient profiles and authentic conversations, which are rigorously assessed by both AI judges and human clinicians to ensure clinical accuracy and naturalness. This resource provides a scalable benchmark for developing and improving automated depression diagnosis systems.
The global demand for mental health services, particularly for conditions like major depressive disorder, far exceeds the availability of qualified professionals and real-world training data. This significant gap limits the development and evaluation of effective diagnostic tools. To address this challenge, researchers have explored the creation of simulated or virtual patients, but existing methods often struggle to produce clinically accurate, natural, and diverse symptom presentations.
A new approach, called TalkDep, offers a promising solution. Developed by researchers including Xi Wang, Anxo Perez, Javier Parapar, and Fabio Crestani, TalkDep is a novel pipeline that leverages advanced large language models (LLMs) to create clinically grounded patient personas for conversation-centric depression screening. The core idea is to simulate authentic patient responses that can better support the training and evaluation of diagnostic models.
How TalkDep Creates Realistic Patient Personas
TalkDep employs a “clinician-in-the-loop” strategy, meaning mental health professionals are deeply involved in the design and validation process. The creation of a simulated patient persona involves two main stages:
First, realistic patient profiles are constructed. A team of clinical psychologists designs templates that summarize biographical and clinical information. These templates include essential attributes such as basic demographic details (name, age, gender, and a predefined depression severity score based on the BDI-II scale), key negative symptoms that the LLM should manifest, a background of personal history and social context for narrative coherence, and a defined communication style that reflects linguistic markers often associated with depression.
Second, the system uses an In-Context Learning (ICL) simulation advancement to generate realistic dialogues. Based on the patient profiles, an initial LLM generates synthetic conversations. For instance, it might create five dialogues: one representing the overall depression severity and four highlighting individual main symptoms. These generated dialogues are then framed as counseling sessions and assessed by another LLM, which acts as a clinical professional to predict the severity of depression. If this assessment aligns closely with the ground truth depression profile (an absolute difference in BDI score of less than five points), the conversations are combined with the patient profile to form a complete context for the final patient simulator. If not, the generation and evaluation process is repeated with refined prompts. This rigorous process has resulted in 12 fully validated personas, covering a range from minimal to severe depression.
Validating the Simulated Patients
The reliability of TalkDep’s simulated patients was verified through extensive assessments by both LLM judges and human clinical professionals.
In one experiment, LLMs were used as judges in a pairwise comparison. They were given two conversation transcripts from different simulated personas and asked to identify which one reflected a higher risk of depression. Among the four open-source models tested, DeepSeek R1-14B achieved the highest accuracy at 86.36%, demonstrating that the depression signals embedded in the dialogues are detectable and interpretable by other LLMs.
To complement this, two certified clinical psychologists, who were not involved in the simulation design, independently rated the 12 LLM personas. They assessed the realism of the simulated patients and their behaviors on a 1-5 Likert scale across two dimensions: general interaction quality (humanness, naturalness, fluency) and depression diagnosis-oriented assessment (emotional consistency, symptom realism, engagement/responsiveness, cognitive load & processing style). The overall mean score across all personas and attributes was 3.92, indicating that clinicians found the simulated patients to be well above the midpoint of adequacy.
Interestingly, while the general quality of interaction slightly increased with severity, the depression-oriented attributes were rated better in the minimal and moderate bands, declining slightly for severe cases. This finding is consistent with clinical expectations, as severely depressed patients often exhibit reduced participation and cognitive slowing.
Also Read:
- Advancing Depression Assessment: A New Dataset and AI Reasoning Approach
- A New Framework for Trustworthy Medical AI Evaluation
Impact and Future Directions
The TalkDep simulation pipeline has proven to produce clinically relevant and conversationally coherent patient personas across various depression severity levels. It has already been utilized as the ground-truth backbone for the eRisk 2025 pilot track, which focuses on conversational depression detection. This resource offers a scalable and adaptable tool for improving the robustness and generalizability of automatic depression diagnosis systems.
The availability of these validated simulated patients provides a solid foundation for broader research in patient simulation, particularly for mental health education. It enables controlled experimentation with symptom severity and multi-system evaluation, paving the way for further advancements in generating and assessing simulated mental health dialogues. It’s important to note that all stages of profile design and evaluation were performed in consultation with licensed clinicians, and no real patient data was used, ensuring ethical compliance and no risk of re-identification or psychological harm to actual individuals. You can learn more about this research by reading the full paper here.


