Qwen3-ASR Atypical English — Encoder Adaptation + SpecAugment
This repository contains the validation-selected encoder checkpoint from the
promptless Qwen3-ASR atypical-English architectural ablation.
Experiment
- Base model:
Qwen/Qwen3-ASR-1.7B
- Dataset:
cdli/ugandan_english_nonstandard_speech_v1.0
- Condition:
encoder
- Prompt: none
- SpecAugment: enabled, train-only
- LR:
5e-05
- Scheduler:
cosine
- Seed:
42
- Batch / grad accumulation:
4 / 4
- Effective batch:
16
SpecAugment
| Setting | Value |
|---|
| Probability | 0.5 |
| Time masks | 2 |
| Time width | 40 |
| Frequency masks | 2 |
| Frequency width | 12 |
Trainable parameters
| Metric | Value |
|---|
| Total | 2,038,052,480 |
| Trainable | 314,328,704 |
| Trainable % | 15.4230% |
Validation checkpoint selection
Selection was performed on the validation split only using
mean normalized per-utterance WER/CER capped at 1.0.
| Metric | Value |
|---|
| Selected checkpoint | checkpoint-1500 |
| Mean capped WER | 31.33% |
| Mean capped CER | 22.29% |
| Corpus WER | 30.88% |
| Corpus CER | 21.17% |
Final test results
Primary reporting uses normalized per-utterance WER/CER, capped at 1.0 per
utterance and then averaged. Normalized corpus metrics are retained as
secondary diagnostics.
| Metric | Result |
|---|
| Mean capped utterance WER | 22.59% |
| Mean capped utterance CER | 14.30% |
| Normalized corpus WER | 23.54% |
| Normalized corpus CER | 14.62% |
| Relative primary WER reduction vs base | 32.45% |
Results by severity
| Group | n | Speakers | Mean capped WER | Mean capped CER | Corpus WER | Corpus CER |
|---|
| Mild (easily understood with minimal effort) | 334 | 3 | 20.21% | 11.76% | 23.12% | 13.09% |
| Moderate (requires effort to understand) | 340 | 3 | 23.23% | 14.32% | 22.37% | 13.79% |
| Severe (frequent breakdowns) | 339 | 3 | 24.30% | 16.78% | 25.05% | 17.11% |
Results by disorder
| Group | n | Speakers | Mean capped WER | Mean capped CER | Corpus WER | Corpus CER |
|---|
| Articulation Disorders | 176 | 2 | 21.69% | 11.72% | 22.79% | 12.19% |
| Dysarthria | 253 | 2 | 23.67% | 14.80% | 23.79% | 14.86% |
| Stuttering (Disfluency Disorders) | 488 | 4 | 21.54% | 14.20% | 22.23% | 14.50% |
| Voice disorder | 96 | 1 | 26.72% | 18.18% | 28.67% | 19.60% |
Reproducibility
The exact corrected trainer is archived in this repository under:
code/qwen3_asr_sft_fixed_ablation.py
The companion private results repository is:
KasuleTrevor/qwen3-asr-en-atypical-promptless-specaug-results
No raw participant audio is redistributed.