RaahNaameh-1 Textual Corpus
A large-scale Persian text corpus assembled for training the RaahNaameh-1 embedding model.
Source
Sentences
Description
Jomleh
1,002,221
Formal Persian web text
LSCP
10,257,866
Iranian tweets — colloquial, slang, emoji
Persian Wikipedia
1,107,618
Encyclopedic articles
Total
12,367,705
Light normalization only: Arabic→Persian character mapping, zero-width space removal
Emojis, Finglish… See the full description on the dataset page:
https://huggingface.co/datasets/Reza2kn/RaahNaameh-1-textual-corpus.