Base model: MohamedRashad/Arabic-Whisper-CodeSwitching-Edition
Baseline pretraining data included MohamedRashad/arabic-english-code-switching
Private Bahraini speech corpus with manual transcript cleanup
Dialect-specific correction passes for Bahraini phrasing
Review-driven checkpoint selection from unseen clips
The public release is packaged under bahraini_asr_codeswitching
Why Bahraini fine-tuning matters
The baseline model is useful, but its upstream code-switching data includes a broader Arabic mix and is not specialized for Bahraini dialect preservation. In practice, that means some outputs can be semantically plausible while still drifting away from the speaker's local wording.
This matters for Bahraini ASR because a transcript can look "close enough" in standard Arabic or another dialect while still failing to preserve the actual form that was spoken.
Bahraini ASR overview
Evaluation summary
This checkpoint was selected from a manual unseen-clip review rather than a formal WER/CER benchmark.
On the original 50-clip checkpoint-1250 review sheet:
ours: 22
tie: 20
baseline: 4
both_bad: 4
Raw counts from that sheet suggest a strong advantage for the fine-tuned model, but for public-facing reporting we should be more conservative than the raw sheet alone:
a reasonable conservative takeaway is that the fine-tuned model was better on at least roughly 60%+ of meaningful cases
it was at minimum competitive with the baseline on a clear majority of the reviewed clips
exact percentages should be treated as approximate because some rows were later re-checked manually and a few public-facing examples were judged more cautiously
This original review result is the main reason checkpoint-1250 was chosen as the public release over later checkpoints.
Important note:
these counts come from the original checkpoint-1250 review CSV
the local example table below was re-vetted manually afterward
some later local proxy review files had labeling inconsistencies, so the example table is treated as the more conservative public-facing evidence set
in other words: the counts below are from the original 1250 review pass, while the example rows were hand-pruned to avoid overstating wins
Dialect preservation candidates
In addition to the manual 50-clip review, we ran a separate text-only mining pass over the private training corpus to look for likely dialect-preservation patterns between:
text
text_asr_v5_raw
This is not a benchmark and should not be read as ASR accuracy. It is a heuristic analysis meant to surface likely cases where a local Bahraini form was shifted toward a more common outside-dialect form.
heuristic dialect-shift candidates account for about 0.72% of changed rows
the stricter local-to-common subset accounts for about 0.07% of changed rows
these are small percentages in absolute terms, but they are high-value because they target the exact dialect-preservation failure mode we care about
Most common strict patterns found:
ليش -> ليه: 19
مب -> مش: 18
Useful framing:
some baseline outputs are semantically plausible, but they do not preserve the speaker's original dialect wording
this analysis is best treated as dialect preservation candidate mining, not formal detection accuracy
Review examples
These clips are the strongest remaining ours wins from the corrected local review file after removing rows that we manually judged as baseline, tie, or both_bad.
Hugging Face model cards do not show a built-in inline playbar for these samples here. Click a .wav link below and Hugging Face/browser will open the file in its audio player.