Significance of Speaker Embeddings and Temporal Context for Depression Detection

doi:10.48550/arXiv.2107.13969

Significance of Speaker Embeddings and Temporal Context for Depression Detection

Depression detection from speech has attracted a lot of attention in recent years. However, the significance of speaker-specific information in depression detection has not yet been explored. In this work, we analyze the significance of speaker embeddings for the task of depression detection from speech. Experimental results show that the speaker embeddings provide important cues to achieve state-of-the-art performance in depression detection. We also show that combining conventional OpenSMILE and COVAREP features, which carry complementary information, with speaker embeddings further improves the depression detection performance. The significance of temporal context in the training of deep learning models for depression detection is also analyzed in this paper.

Publication:

arXiv e-prints

Pub Date:

July 2021

DOI:

10.48550/arXiv.2107.13969

arXiv:

arXiv:2107.13969

Bibcode:

2021arXiv210713969H

Keywords:

Computer Science - Computers and Society;
Computer Science - Machine Learning;
Computer Science - Sound;
Electrical Engineering and Systems Science - Audio and Speech Processing

NASA/ADS

Significance of Speaker Embeddings and Temporal Context for Depression Detection

Abstract