This project provides a Python-based pipeline to load and access pre-processed legal case datasets stored in the Hugging Face datasets format. The datasets are intended for Natural Language Processing (NLP) tasks such as token classification, masked language modeling, and span extraction.
The datasets are stored in Arrow format using Hugging Face's datasets library. There are four main datasets:
Contains masked versions… See the full description on the dataset page:
https://huggingface.co/datasets/Sanjithganesp/LegalBiasDataset.