A Persian-language instruction-tuning dataset of ~120,000 samples, generated from the
Persian Wikipedia (fawiki) article dump. Each article was cleaned to Markdown, chunked by
section, and passed to a locally-run Gemma4 model that produced grounded instruction/response
pairs across five task types. The result is ready for supervised fine-tuning (SFT) of
Persian LLMs.
Heads-up: this is synthetic data. The instructions and answers were written by an
LLM… See the full description on the dataset page:
https://huggingface.co/datasets/Jamalianpour/persian-wikipedia-instruct.