A prepared Persian article dataset used to train
Haman Persian Article Graph-LLM 125M.
The dataset is derived from the Persian Wikipedia pages-articles dump and was
cleaned and prepared using the native article workflow in the
Rakhshai Graph-based NLP project.
Language: Persian
Accepted articles: 185,906
Training records: 176,611
Validation records: 9,295
Validation ratio: 5%
Split seed: 42
Training format:… See the full description on the dataset page:
https://huggingface.co/datasets/aria-haman/haman-fa-wikipedia-articles-186k.