Harnessing multimodal approaches for depression detection using large language models and facial expressions

Misha Sadeghi; Robert Richer; Bernhard Egger; Lena Schindler-Gmelch; Lydia Helene Rupp; Farnaz Rahimi; Matthias Berking; Bjoern M Eskofier

doi:10.1038/s44184-024-00112-8

Harnessing multimodal approaches for depression detection using large language models and facial expressions

Npj Ment Health Res. 2024 Dec 23;3(1):66. doi: 10.1038/s44184-024-00112-8.

Authors

Misha Sadeghi¹, Robert Richer², Bernhard Egger³, Lena Schindler-Gmelch⁴, Lydia Helene Rupp⁴, Farnaz Rahimi², Matthias Berking⁴, Bjoern M Eskofier^{2

5}

Affiliations

¹ Machine Learning and Data Analytics Lab (MaD Lab), Department Artificial Intelligence in Biomedical Engineering (AIBE), Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU), Erlangen, 91052, Germany. [email protected].
² Machine Learning and Data Analytics Lab (MaD Lab), Department Artificial Intelligence in Biomedical Engineering (AIBE), Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU), Erlangen, 91052, Germany.
³ Chair of Visual Computing (LGDV), Department of Computer Science, Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU), Erlangen, 91058, Germany.
⁴ Chair of Clinical Psychology and Psychotherapy (KliPs), Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU), Erlangen, 91052, Germany.
⁵ Translational Digital Health Group, Institute of AI for Health, Helmholtz Zentrum München - German Research Center for Environmental Health, Neuherberg, 85764, Germany.

Abstract

Detecting depression is a critical component of mental health diagnosis, and accurate assessment is essential for effective treatment. This study introduces a novel, fully automated approach to predicting depression severity using the E-DAIC dataset. We employ Large Language Models (LLMs) to extract depression-related indicators from interview transcripts, utilizing the Patient Health Questionnaire-8 (PHQ-8) score to train the prediction model. Additionally, facial data extracted from video frames is integrated with textual data to create a multimodal model for depression severity prediction. We evaluate three approaches: text-based features, facial features, and a combination of both. Our findings show the best results are achieved by enhancing text data with speech quality assessment, with a mean absolute error of 2.85 and root mean square error of 4.02. This study underscores the potential of automated depression detection, showing text-only models as robust and effective while paving the way for multimodal analysis.