Text corpora are essential for training models used in tasks like summarization, translation, and large language models (LLMs).
While various efforts have been made to collect monolingual and multilingual datasets in many languages,
Persian has often been underrepresented due to limited resources for data collection and preprocessing.
Existing Persian datasets are typically small and lack content diversity, consisting mainly of… See the full description on the dataset page:
https://huggingface.co/datasets/MatinaAI/matina_persian_text_corpus.