This dataset contains a collection of 16,052,878 unique Arabic words. These words were extracted from a large corpus of Arabic text originating from two primary sources: the Shamela library and the Hindawi library.
Key Characteristics:
Unique Words: The dataset is focused on uniqueness. Each entry in the dataset represents a distinct Arabic word, and duplicates have been removed.
Diacritic Sensitivity: Words with different diacritical markings are considered… See the full description on the dataset page:
https://huggingface.co/datasets/ImruQays/16-million-raw-arabic-words.