Modular Speech-to-Text Translation for Zero-Shot Cross-Modal Transfer
Abstract
Recent research has shown that independently trained encoders and decoders, combined through a shared fixed-size representation, can achieve competitive performance in speech-to-text translation. In this work, we show that this type of approach can be further improved with multilingual training. We observe significant improvements in zero-shot cross-modal speech translation, even outperforming a supervised approach based on XLSR for several languages.
- Publication:
-
arXiv e-prints
- Pub Date:
- October 2023
- DOI:
- 10.48550/arXiv.2310.03724
- arXiv:
- arXiv:2310.03724
- Bibcode:
- 2023arXiv231003724D
- Keywords:
-
- Computer Science - Computation and Language
- E-Print:
- Proceedings of Interspeech 2023