Classifying Emotion Expression in Mental Health Statements: A Comparison of Professionals, Laypeople, and LLMs
Le résumé fourni par la source
This study investigates inter-rater agreement and classification tendencies when identifying emotion expression in mental health context statements. Using a secondary dataset of 100 unique Reddit statements expressing emotional distress, we evaluate annotations across three distinct rater populations: Mental Health Professionals, Laypeople, and Large Language Models (GPT-3.5-turbo and GPT-4o-mini). First, we test whether mental health professionals demonstrate higher internal consensus (Krippendorff’s alpha) compared to laypeople when evaluating emotion expression. Second, we use Binary Logistic Mixed-Effects Models to examine whether overall classification tendencies differ between professionals and laypeople. Third, we test whether zero-shot automated ratings from LLMs align more closely with the rating tendencies of mental health professionals or those of laypeople. As exploratory extensions, we evaluate whether custom prompt engineering (via pre-annotated data from Safespace Research) improves LLM alignment with mental health professionals compared to zero-shot prompting, and examine whether specific statement-level metadata (e.g., word count) predicts classification disagreement across groups.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.