Byakto Speech: Real-time long speech synthesis with convolutional neural network: Transfer learning from English to Bangla

doi:10.48550/arXiv.2106.03937

Byakto Speech: Real-time long speech synthesis with convolutional neural network: Transfer learning from English to Bangla

Speech synthesis is one of the challenging tasks to automate by deep learning, also being a low-resource language there are very few attempts at Bangla speech synthesis. Most of the existing works can't work with anything other than simple Bangla characters script, very short sentences, etc. This work attempts to solve these problems by introducing Byakta, the first-ever open-source deep learning-based bilingual (Bangla and English) text to a speech synthesis system. A speech recognition model-based automated scoring metric was also proposed to evaluate the performance of a TTS model. We also introduce a test benchmark dataset for Bangla speech synthesis models for evaluating speech quality. The TTS is available at https://github.com/zabir-nabil/bangla-tts

Publication:

arXiv e-prints

Pub Date:

May 2021

DOI:

10.48550/arXiv.2106.03937

arXiv:

arXiv:2106.03937

Bibcode:

2021arXiv210603937A

Keywords:

Computer Science - Sound;
Computer Science - Machine Learning;
Electrical Engineering and Systems Science - Audio and Speech Processing

NASA/ADS

Byakto Speech: Real-time long speech synthesis with convolutional neural network: Transfer learning from English to Bangla

Abstract