adapters library, specifically targeting the first 4 layers of the Encoder (for acoustic/accent adaptation) and the full Decoder (for medical jargon and linguistic structure).bnb.optim.AdamW8bit)| Training Loss | Epoch | Step | Validation Loss | WER (%) |
|---|---|---|---|---|
| No log | 3.03 | 100 | 0.0778 | 133.333 |
| No log | 6.06 | 200 | 0.0579 | 14.2827 |
| No log | 9.09 | 300 | 0.0542 | 15.6211 |
| No log | 12.12 | 400 | 0.0514 | 7.42367 |
| 0.0761 | 15.15 | 500 | 0.0465 | 8.42744 |
| No log | 18.18 | 600 | 0.0450 | 6.41991 |
| No log | 21.21 | 700 | 0.0457 | 6.44082 |
| No log | 24.24 | 800 | 0.0458 | 6.37808 |
| No log | 27.27 | 900 | 0.0458 | 6.29444 |
| 0.0003 | 30.30 | 1000 | 0.0464 | 8.26014 |
| No log | 33.33 | 1100 | 0.0466 | 8.30197 |
| No log | 36.36 | 1200 | 0.0466 | 8.23923 |
| No log | 39.39 | 1300 | 0.0468 | 8.19741 |
| 0.0001 | 42.42 | 1400 | 0.0468 | 8.19741 |
temperature = 0.0) across all models for a fair comparison.| Rank | Model | WER (%) | CER (%) | Sentence Accuracy (%) |
|---|---|---|---|---|
| 1 | Whisper-AfroRad-FR (this model) | 20.93 | 16.80 | 34.67 |
| 2 | Med-Whisper-AfroRad-FR | 21.84 | 17.68 | 29.33 |
| 3 | whisper-small-rad-FR | 25.12 | 20.89 | 33.33 |
| 4 | nvidia/canary-1b-v2 | 33.96 | 11.10 | 1.33 |
| 5 | Qwen/Qwen3-ASR-0.6B | 45.40 | 17.55 | 0.00 |
| 6 | bofenghuang/whisper-small-cv11-french | 75.11 | 53.65 | 0.00 |
| 7 | openai/whisper-small (baseline) | 79.12 | 54.47 | 0.00 |
| 8 | openai/whisper-large-v3 | 120.41 | 84.02 | 0.00 |
@misc{whisper-afrorad-fr,
author = {StephaneBah},
title = {Whisper-AfroRad-FR: Medical Radiology ASR for Afro-French Context},
year = {2026},
publisher = {Hugging Face},
howpublished = {\\url{https://huggingface.co/StephaneBah/Whisper-AfroRad-FR}}
}