Whisper Large V3 — Luganda Stage 1 (Waxal, no SpecAugment)
Stage 1 fine-tune of openai/whisper-large-v3 on the WaxalNLP Luganda standard speech dataset.
This is the original Stage 1 checkpoint without SpecAugment, used as the base for early Stage 2 runs (v16, v17, v20).
Training Details
Dataset: google/WaxalNLP (lug_asr split)
Language token: sw (Swahili used as proxy for Luganda)
Full model training: encoder, decoder, and projection updated