HmBlogs is a general Persian corpus collected from nearly 20 million blog posts over a period of 15 years containig 6.8 billion tokens.
This version is the preprocessed version of the dataset prepared by the original authors and converted to proper format to integrate with 🤗Datasets.
In order to access the raw versions visit the official link at
http://nlplab.sbu.ac.ir/hmBlogs-v3 .
Paper:
https://arxiv.org/abs/2111.02362
Authors: Hamzeh… See the full description on the dataset page:
https://huggingface.co/datasets/sbunlp/hmblogs-v3.