This checkpoint was trained on the full 960-hour LibriSpeech training set. During training, the MMS-300M speech encoder and Q-Former were trainable, while the Aya decoder remained frozen.
Evaluation
The model was evaluated on the LibriSpeech test-clean split.