The Sindhi Legal Dataset is a parallel text dataset designed for legal NLP tasks in low-resource Sindhi language settings. It contains Pakistani legal QnA in Sindhi (Arabic script).
The dataset is intended for applications such as legal text generation, machine translation, retrieval-augmented generation (RAG), and legal question answering systems.
Each record consists of an input legal text segment and its corresponding output legal text segment.… See the full description on the dataset page:
https://huggingface.co/datasets/DanishMahdi/Sindhi_Legal_1.