Modular Speech-to-Text Translation for Zero-Shot Cross-Modal Transfer

doi:10.48550/arXiv.2310.03724

Modular Speech-to-Text Translation for Zero-Shot Cross-Modal Transfer

Recent research has shown that independently trained encoders and decoders, combined through a shared fixed-size representation, can achieve competitive performance in speech-to-text translation. In this work, we show that this type of approach can be further improved with multilingual training. We observe significant improvements in zero-shot cross-modal speech translation, even outperforming a supervised approach based on XLSR for several languages.

Publication:

arXiv e-prints

Pub Date:

October 2023

DOI:

10.48550/arXiv.2310.03724

arXiv:

arXiv:2310.03724

Bibcode:

2023arXiv231003724D

Keywords:

Computer Science - Computation and Language

E-Print:

Proceedings of Interspeech 2023

NASA/ADS

Modular Speech-to-Text Translation for Zero-Shot Cross-Modal Transfer

Abstract