The Sindhi Gold Corpus is a highly curated, mathematically balanced, and rigorously sanitized dataset containing 1.67 GB of pure
Sindhi text. Comprising exactly 5,424,753 continuous lines, this corpus was specifically engineered for continued pre-training and
Masked Language Modeling (MLM) of transformer architectures (such as mBERT, ELECTRA, and DistilBERT).This dataset is a surgically… See the full description on the dataset page:
https://huggingface.co/datasets/Kashif786/sindhi-gold-corpus-mlm.