Noise Tokens: Learning Neural Noise Templates for Environment-Aware Speech Enhancement

doi:10.48550/arXiv.2004.04001

Noise Tokens: Learning Neural Noise Templates for Environment-Aware Speech Enhancement

In recent years, speech enhancement (SE) has achieved impressive progress with the success of deep neural networks (DNNs). However, the DNN approach usually fails to generalize well to unseen environmental noise that is not included in the training. To address this problem, we propose "noise tokens" (NTs), which are a set of neural noise templates that are jointly trained with the SE system. NTs dynamically capture the environment variability and thus enable the DNN model to handle various environments to produce STFT magnitude with higher quality. Experimental results show that using NTs is an effective strategy that consistently improves the generalization ability of SE systems across different DNN architectures. Furthermore, we investigate applying a state-of-the-art neural vocoder to generate waveform instead of traditional inverse STFT (ISTFT). Subjective listening tests show the residual noise can be significantly suppressed through mel-spectrogram correction and vocoder-based waveform synthesis.

Publication:

arXiv e-prints

Pub Date:

April 2020

DOI:

10.48550/arXiv.2004.04001

arXiv:

arXiv:2004.04001

Bibcode:

2020arXiv200404001L

Keywords:

Electrical Engineering and Systems Science - Audio and Speech Processing

E-Print:

5 pages, Submitted to Interspeech 2020

NASA/ADS

Noise Tokens: Learning Neural Noise Templates for Environment-Aware Speech Enhancement

Abstract