Views
No views yet
multi (single≈AbdelAzizErwi), moderate text
normalization, audio-quality filtering (SNR≥8.0dB), phonemization=False,
30 epochs. Targets pronunciation/vowel/consonant quality by improving the signal
rather than the architecture. Deploy raw or EMA weights per the A/B in §8.model.pt + vocab.txt + config.yaml.
The vocab.txt here matches this run (SILMA char vocab, or phonemized vocab if piloted).1@misc{linagora2024Linto-tn,
2 title={LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect},
3 author={Hedi Naouara and Jerome Louradour and Jean-Pierre Lorre},
4 year={2025}, eprint={2504.02604}, archivePrefix={arXiv}, primaryClass={cs.CL}
5}