Fine-tuned checkpoint of
OmniLingual ASR CTC 300M_v2
on a mixture of FLEURS Chichewa/Nyanja (2,653 examples, ~10.87 hours) and the Chichewa Trigrams
speech-text parallel corpus (25,991 examples, ~10.87 hours — 100% of FLEURS train hours).
Training used the fairseq2 wav2vec2 ASR recipe for 5,000 steps with the encoder frozen for the
first 1,000 steps. Tokenizer:
omniASR_tokenizer_written_v2.
This model outperforms the zero-shot omniASR CTC 3B baseline (35.80% WER on FLEURS,
62.10% on Zambezi Voice) despite having 10× fewer parameters. Compared to fine-tuning
on FLEURS alone (36.53% FLEURS / 61.93% Zambezi), the trigram augmentation reduces
WER by 1.26 pp on FLEURS and 1.14 pp on Zambezi.
The checkpoint is stored in fairseq2's sharded format.
checkpoint/model/pp_00/tp_00/sdp_00.pt contains the full model state
(training used a single GPU, so there is only one shard).