spot_img
HomeResearch & DevelopmentCARMA: A New Dataset for Understanding Arabic Mental Health...

CARMA: A New Dataset for Understanding Arabic Mental Health on Reddit

TLDR: CARMA is the first large-scale, automatically annotated dataset of Arabic Reddit posts designed for mental health research. It covers six mental health conditions (ADHD, Anxiety, Autism, Depression, OCD, Suicide) and a control group, addressing the critical scarcity of Arabic-language mental health data. By analyzing linguistic patterns, CARMA enables the development of more effective detection models, offering a valuable resource for advancing mental health support in underrepresented languages.

Mental health disorders are a global concern, impacting millions worldwide. However, early detection and support remain significant challenges, particularly within Arabic-speaking populations. This is often due to cultural stigma surrounding mental health and a scarcity of specialized resources. While extensive research has focused on English-language mental health detection, the Arabic language has remained largely unexplored in this critical area, primarily because of a lack of annotated datasets.

Addressing this crucial gap, researchers have introduced CARMA, the first automatically annotated, large-scale dataset of Arabic Reddit posts. This groundbreaking dataset covers six distinct mental health conditions, including Anxiety, Autism, ADHD, Depression, OCD, and Suicide, alongside a comprehensive control group. CARMA significantly surpasses existing resources in both its sheer scale and the diversity of conditions it encompasses.

The creation of CARMA involved a meticulous process. Researchers collected data from general Arabic subreddits, as dedicated mental health forums are less common in Arabic compared to English. They then identified users who self-reported a diagnosis using carefully curated diagnosis patterns and keywords. This method, inspired by successful English-language studies, allowed for automatic annotation, making the dataset scalable. To ensure data quality, a rigorous cleaning pipeline was applied, filtering out non-Arabic content and short posts, and excluding any control users who showed signs of mental health conditions.

One of CARMA’s key contributions is its in-depth qualitative and quantitative analysis of lexical and semantic differences between users. This analysis provides invaluable insights into the unique linguistic markers associated with specific mental health conditions in Arabic discourse. For instance, Anxiety-related posts often feature first-person and introspective language, while Depression posts show strong self-disclosure. Autism-related language tends to be more detached and analytical, reflecting discussions about interpersonal experiences rather than emotional sharing.

To demonstrate the dataset’s potential, classification experiments were conducted using a range of models, from traditional classifiers to advanced large language models. The results highlight the significant promise of CARMA in advancing mental health detection in underrepresented languages like Arabic. For example, an SVM classifier achieved an F1 score of 0.83 for Anxiety detection, showcasing the effectiveness of Arabic-specific transformers.

CARMA is not just a dataset; it’s a foundational resource for future research. It lays the groundwork for developing predictive tools, enabling earlier detection and intervention, and fostering a deeper understanding of mental health discourse within Arabic-speaking communities. Researchers can access the dataset and learn more about its methodology through its GitHub repository and Hugging Face page.

Also Read:

This initiative marks a significant step towards more inclusive and data-driven mental health research, ensuring that the unique linguistic and cultural expressions of mental distress in Arabic are recognized and addressed.

Rhea Bhattacharya
Rhea Bhattacharyahttps://blogs.edgentiq.com
Rhea Bhattacharya is an AI correspondent with a keen eye for cultural, social, and ethical trends in Generative AI. With a background in sociology and digital ethics, she delivers high-context stories that explore the intersection of AI with everyday lives, governance, and global equity. Her news coverage is analytical, human-centric, and always ahead of the curve. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -