Qwen3-ASR Atypical Luganda — Projector Adaptation + SpecAugment
This repository contains the validation-selected projector checkpoint from the
promptless Qwen3-ASR atypical-Luganda architectural ablation.
Experiment
- Base model:
KasuleTrevor/cdli-qwen3-asr-lg-typical-1p7b-base-finetune
- Dataset:
cdli/ugandan_luganda_nonstandard_speech_v1.0
- Condition:
projector
- Prompt: none
- SpecAugment: enabled, train-only
- LR:
0.0001
- Scheduler:
cosine
- Seed:
42
- Batch / grad accumulation:
4 / 4
- Effective batch:
16
SpecAugment
| Setting | Value |
|---|
| Probability | 0.5 |
| Time masks | 2 |
| Time width | 40 |
| Frequency masks | 2 |
| Frequency width | 12 |
Trainable parameters
| Metric | Value |
|---|
| Total | 2,038,052,480 |
| Trainable | 3,148,800 |
| Trainable % | 0.1545% |
Validation checkpoint selection
Selection was performed on the validation split only using
normalized corpus WER/CER.
| Metric | Value |
|---|
| Selected checkpoint | checkpoint-1500 |
| Mean capped WER | 72.85% |
| Mean capped CER | 39.79% |
| Corpus WER | 89.93% |
| Corpus CER | 48.16% |
Final test results
Primary reporting uses normalized per-utterance WER/CER, capped at 1.0 per
utterance and then averaged. Normalized corpus metrics are retained as
secondary diagnostics.
| Metric | Result |
|---|
| Mean capped utterance WER | 69.33% |
| Mean capped utterance CER | 33.19% |
| Normalized corpus WER | 83.62% |
| Normalized corpus CER | 43.83% |
| Relative primary WER reduction vs base | 9.57% |
Results by severity
| Group | n | Speakers | Mean capped WER | Mean capped CER | Corpus WER | Corpus CER |
|---|
| Mild (easily understood with minimal effort) | 366 | 3 | 62.10% | 25.64% | 71.93% | 33.40% |
| Moderate (requires effort to understand) | 347 | 3 | 67.13% | 29.85% | 73.99% | 33.34% |
| Severe (frequent breakdowns) | 315 | 3 | 80.16% | 45.65% | 108.48% | 70.48% |
Results by disorder
| Group | n | Speakers | Mean capped WER | Mean capped CER | Corpus WER | Corpus CER |
|---|
| Articulation Disorders | 209 | 2 | 67.18% | 27.06% | 76.76% | 36.46% |
| Dysarthria | 276 | 2 | 59.94% | 24.62% | 65.86% | 28.67% |
| Stuttering (Disfluency Disorders) | 438 | 4 | 72.81% | 37.22% | 82.89% | 41.81% |
| Voice disorder | 105 | 1 | 83.78% | 51.12% | 129.59% | 88.20% |
Reproducibility
The exact corrected trainer is archived in this repository under:
code/qwen3_asr_sft_fixed_ablation.py
The companion private results repository is:
KasuleTrevor/qwen3-asr-lg-atypical-promptless-specaug-results
No raw participant audio is redistributed.