DunbaaBERT is a family of Urdu RoBERTa-base encoder models trained from scratch on a deduplicated 17 GB Urdu corpus. The models use Byte-BPE vocabularies of 32k, 52k, and 96k tokens and are released under the MIT license.
The final corpus was constructed from multiple Urdu resources and deduplicated at line level.
We report a normalized efficiency metric combining Macro-F1 and inference throughput.
Across benchmarks, the DunbaaBERT family consistently achieved stronger performance-efficiency trade-offs than most multilingual baselines.
DunbaaBERT-52k achieved the strongest linguistic probing performance on UrBLiMP, while DunbaaBERT-32k provided the strongest overall efficiency profile.
Interestingly, DunbaaBERT-96k ranked second in average efficiency despite having the largest vocabulary.
Get the fairseq checkpoint
here.
1@misc{maab2026dunbaabertsacrificesemantics,
2 title={DunbaaBERT: From Sacrifice to Semantics},
3 author={Iffat Maab and Waleed Jamil and Raphael Schmitt},
4 year={2026},
5 eprint={2605.26935},
6 archivePrefix={arXiv},
7 primaryClass={cs.CL},
8 url={https://arxiv.org/abs/2605.26935},
9}