This is a highly curated, deduplicated, and linguistically sanitized conversational dataset in Pashto. It is specifically structured for instruction tuning, alignment, and perfecting the conversational capabilities of local Large Language Models (LLMs) in the Pashto language.
Total Conversations: 722 unique multi-turn/single-turn interactions (derived from 907 raw instances).
Format: OpenAI / Hugging Face… See the full description on the dataset page:
https://huggingface.co/datasets/nassimjp/pashto-clean-conversations.