This repository contains the Hadith Alpaca Dataset, comprising 4,592 meticulously processed Hadiths.
The dataset removes the initial chain of transmission in Arabic and extraneous commentary, focusing on the core Hadith text.
It's designed for training and evaluating language models, particularly in understanding and processing Islamic religious texts.