Fleurs Synthetic Whisper Medium Fleurs Synthetic Hours0P75 LoRA Adapter
Summary
This repository contains a Whisper checkpoint for Chichewa/Nyanja
automatic speech recognition, fine-tuned from openai/whisper-medium.
- Experiment type:
fleurs-synthetic
- Base model:
openai/whisper-medium
- Training condition:
fleurs_synthetic_hours0p75
- Release artifact: LoRA adapter checkpoint selected from the best training checkpoint
Intended use
This adapter is intended for research and evaluation on Chichewa/Nyanja ASR. It must be used together with the base Whisper model.
It is not a production-ready speech system and should be validated carefully
before downstream use.
How to use
This repository contains an adapter, not a fully merged standalone model. Load
the base model first, then attach the adapter.
1from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor
2from peft import PeftModel
3
4base_model_id = "openai/whisper-medium"
5adapter_repo_id = "ai4good-labyrinth/fleurs-synthetic-hours0p75-whisper-medium-no-language-lora-adapter"
6
7processor = AutoProcessor.from_pretrained(base_model_id)
8base_model = AutoModelForSpeechSeq2Seq.from_pretrained(base_model_id)
9model = PeftModel.from_pretrained(base_model, adapter_repo_id)
The local evaluation script in this repository can also load the adapter
directly because it reads adapter_config.json and automatically fetches the
base model.
Training data
- Training source:
FLEURS train + Prepared dataset
- Evaluation source during training:
FLEURS dev
- Train examples before duration filtering:
8243
- Train examples after duration filtering:
8202
- Dev examples before duration filtering:
311
- Dev examples after duration filtering:
305
- Duration filter used during training:
min_duration_seconds=0.0, max_duration_seconds=30.0
Training procedure
- Fine-tuning script:
experiments/whisper_finetune/finetune_whisper.py
- Base model:
openai/whisper-medium
- Task:
transcribe
- Language hint during training/evaluation: none, corresponding to
--language auto in standalone evaluation
- LoRA:
yes
- LoRA rank:
32
- LoRA alpha:
64
- LoRA dropout:
0.05
- LoRA target modules:
fc1,out_proj,k_proj,q_proj,fc2,v_proj
- Extra trainable modules:
embed_tokens,proj_out
- Mixed precision:
fp16
- Gradient checkpointing:
True
- Selected checkpoint step:
2600
- Selected checkpoint epoch:
2.53
Training-time dev selection
The best checkpoint was selected using trainer-side dev evaluation on the
duration-filtered FLEURS dev split.
- Dev WER:
0.4847
- Dev CER:
0.1220
- Dev loss:
0.6190
These values come from the training pipeline and may differ slightly from
standalone post-hoc evaluation because the decoding path is not perfectly
identical.
Evaluation protocol
Standalone evaluation is recommended for the final release. Filtered and
unfiltered results should be reported separately.
- Filtered evaluation:
min_duration_seconds=0, max_duration_seconds=30
- Unfiltered evaluation: no duration constraint
- Decoding task:
transcribe
- Language hint:
auto
Evaluation summary
| Dataset | Split | Setting | Num examples | WER | CER | Notes |
|---|
| FLEURS | dev | filtered | 305 | 0.5378 | 0.1363 | Filtered to 30 seconds |
| FLEURS | dev | unfiltered | TBD | TBD | TBD | Standalone eval pending |
| FLEURS | test | filtered | 745 | 0.5018 | 0.1354 | Filtered to 30 seconds |
| FLEURS | test | unfiltered | TBD | TBD | TBD | Standalone eval pending |
| Zambezi | dev | filtered | 613 | 0.6410 | 0.1818 | Filtered to 30 seconds |
| Zambezi | dev | unfiltered | TBD | TBD | TBD | Standalone eval pending |
| Zambezi | test | filtered | 427 | 0.6322 | 0.1489 | Filtered to 30 seconds |
| Zambezi | test | unfiltered | TBD | TBD | TBD | Standalone eval pending |
Files in this repository
- Adapter weights and config: repository root
- Processor/tokenizer files: repository root
- Evaluation JSON files:
eval/...
Known limitations
- Whisper does not provide an official Nyanja/Chichewa language token.
- This repository contains an adapter only, so users must also comply with the upstream base model license and dataset licenses.
- Standalone evaluation and trainer-side evaluation can differ slightly even on the same split and duration filter.
- Cross-dataset results should be interpreted carefully because transcription conventions may differ across corpora.
Citation
If you use this checkpoint, please cite:
- the Whisper paper
- the FLEURS dataset
- this repository
1@misc{fleurs_synthetic_hours0p75_whisper_medium_no_language_lora_adapter_2026,
2 title = {Fleurs Synthetic Whisper Medium Fleurs Synthetic Hours0P75 LoRA Adapter},
3 author = {AI4Good Labyrinth Team},
4 year = {2026},
5 howpublished = {\url{https://huggingface.co/ai4good-labyrinth/fleurs-synthetic-hours0p75-whisper-medium-no-language-lora-adapter}},
6 note = {Whisper LoRA adapter fine-tuning for Chichewa/Nyanja ASR}
7}