A continuously growing high-quality Pashto text corpus designed for Continual Pretraining (CPT) of large language models.
The corpus is intended to provide broad linguistic coverage of modern written Pashto while preserving natural text. It is suitable for language modeling, tokenizer development, vocabulary analysis, and large-scale pretraining.
Native Pashto text
UTF-8 encoded
Plain text format
One document per line
Duplicate… See the full description on the dataset page:
https://huggingface.co/datasets/nassimjp/pashto-cpt-corpus.