AutoMode-ASR: Learning to Select ASR Systems for Better Quality and Cost

doi:10.48550/arXiv.2409.12476

AutoMode-ASR: Learning to Select ASR Systems for Better Quality and Cost

We present AutoMode-ASR, a novel framework that effectively integrates multiple ASR systems to enhance the overall transcription quality while optimizing cost. The idea is to train a decision model to select the optimal ASR system for each segment based solely on the audio input before running the systems. We achieve this by ensembling binary classifiers determining the preference between two systems. These classifiers are equipped with various features, such as audio embeddings, quality estimation, and signal properties. Additionally, we demonstrate how using a quality estimator can further improve performance with minimal cost increase. Experimental results show a relative reduction in WER of 16.2%, a cost saving of 65%, and a speed improvement of 75%, compared to using a single-best model for all segments. Our framework is compatible with commercial and open-source black-box ASR systems as it does not require changes in model codes.

Publication:

arXiv e-prints

Pub Date:

September 2024

DOI:

10.48550/arXiv.2409.12476

arXiv:

arXiv:2409.12476

Bibcode:

2024arXiv240912476G

Keywords:

Computer Science - Computation and Language;
Computer Science - Sound;
Electrical Engineering and Systems Science - Audio and Speech Processing

E-Print:

SPECOM 2024 Conference

ADS

AutoMode-ASR: Learning to Select ASR Systems for Better Quality and Cost

Abstract