Sanadset 650K: Data on Hadith Narrators
Dataset Description
Sanadset is a large-scale dataset containing over 650,986 Hadith records collected from 926 historical Arabic books. This dataset was created to assist in the computational analysis of Islamic Hadiths, specifically focusing on the chain of narrators (Sanad) and the content (Matn).
It allows researchers to apply Machine Learning and NLP techniques to tasks such as: