This dataset contains a curated, deduplicated collection of high-quality, contextually accurate Pashto linguistic examples. It maps structural language tasks directly to the most critical vocabulary tokens in Pashto, providing a reliable corpus for token warmup, instruction tuning, evaluation, and post-OCR text correction workflows.
The initial release consists of 4,087 verified entries targeting high-frequency and… See the full description on the dataset page:
https://huggingface.co/datasets/nassimjp/pashto-warmup-tokens.