Quran Tajweed ASR — Nemotron Dual-Head v5
Fine-tuned NVIDIA Nemotron 3.5 streaming ASR (0.6B) for full-Quran recitation recognition
with Uthmani script incl. tajweed marks/tashkeel, plus an auxiliary phoneme (tj1) head for
mistake detection.
Best model
- Checkpoint: step 78000 (
quran-nemotron-dual-v5-best-step-78000.nemo, this repo root)
- Held-out canary (128 clips, speaker-disjoint): CER 2.76%, WER 4.43%,
diacritic F1 98.27%, phoneme error rate 3.95%
- Trained: RNNT 0.75 + phoneme-CTC 0.25 dual loss, bf16, cosine schedule,
compact character-transplant vocabulary init (68 Uthmani tokens), all 621M params trainable.
- Data:
tamm5y5m5/quran_with_time — 179,376 train clips (~868 h), speaker-disjoint val/test.
Evaluation sweep (held-out canary)
| Checkpoint | CER | WER | Diacritic F1 | Phoneme ER |
|---|
| 46000 | 3.35% | 5.19% | 97.92% | 4.06% |
| 60000 | 3.26% | 5.19% | 97.97% | 4.02% |
| 63000 | 3.07% | 4.85% | 98.10% | 3.83% |
| 64000 | 3.05% | 4.78% | 98.12% | 3.81% |
| 65000 | 3.06% | 4.85% | 98.10% | 3.92% |
| 66000 | 3.07% | 4.78% | 98.12% | 3.83% |
| 67000 | 3.04% | 4.71% | 98.13% | 3.88% |
| 68000 | 3.03% | 4.71% | 98.15% | 3.93% |
| 69000 | 3.05% | 4.78% | 98.12% | 3.73% |
| 70000 | 2.81% | 4.64% | 98.25% | 3.76% |
| 71000 | 3.22% | 4.92% | 98.04% | 3.83% |
| 72000 | 3.02% | 4.50% | 98.14% | 3.97% |
| 73000 | 3.02% | 4.57% | 98.15% | 3.95% |
| 74000 | 3.00% | 4.57% | 98.13% | 3.86% |
| 75000 | 3.02% | 4.64% | 98.15% | 3.85% |
| 76000 | 3.03% | 4.71% | 98.13% | 3.85% |
| 78000 (best) | 2.76% | 4.43% | 98.27% | 3.95% |
| 78130 | 3.00% | 4.57% | 98.14% | 3.74% |
Note: checkpoints 47k–55k were evaluated during a previous session at similar quality (~3.3% CER);
step selection was made on the lowest-CER checkpoint above.
Inference notes
- Decode with attention context
[56, 3] and the ar-AR language prompt (see stages/full/eval-step-00078000/).
- All training checkpoints (23000–78130, every 1000 steps) are mirrored in dataset repo
tamm5y5m5/quran_nemotron_dual_v5_npl under outputs/nemotron35-quran-dual-v5/full/checkpoints/.
- Known limitation: one out-of-distribution external basmala recording (everyayah-style mastering)
produced an empty hypothesis while held-out data transcribes near-perfectly; base NVIDIA model reads it fine.
Training result (final)
- Target reached: 78,130 steps = 10 epochs; status
complete; min train loss 1.535.