Self-Supervised Attention Networks and Uncertainty Loss Weighting for Multi-Task Emotion Recognition on Vocal Bursts

doi:10.48550/arXiv.2209.07384

Self-Supervised Attention Networks and Uncertainty Loss Weighting for Multi-Task Emotion Recognition on Vocal Bursts

Vocal bursts play an important role in communicating affect, making them valuable for improving speech emotion recognition. Here, we present our approach for classifying vocal bursts and predicting their emotional significance in the ACII Affective Vocal Burst Workshop & Challenge 2022 (A-VB). We use a large self-supervised audio model as shared feature extractor and compare multiple architectures built on classifier chains and attention networks, combined with uncertainty loss weighting strategies. Our approach surpasses the challenge baseline by a wide margin on all four tasks.

Publication:

arXiv e-prints

Pub Date:

September 2022

DOI:

10.48550/arXiv.2209.07384

arXiv:

arXiv:2209.07384

Bibcode:

2022arXiv220907384K

Keywords:

Computer Science - Sound;
Computer Science - Artificial Intelligence;
Electrical Engineering and Systems Science - Audio and Speech Processing

E-Print:

4 pages, 1 figure, accepted at The 2022 ACII Affective Vocal Burst Workshop &amp

NASA/ADS

Self-Supervised Attention Networks and Uncertainty Loss Weighting for Multi-Task Emotion Recognition on Vocal Bursts

Abstract