SafeC4: C4 Dataset with Harmfulness Predictions
Overview
SafeC4 is a processed dataset of the C4 dataset (Colossal, Cleaned version of Common Crawl's web crawl corpus) that includes harmfulness predictions from a HarmFormer As used in our paper - Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs. This dataset can be used for content moderation, safer language model training, or research into harmfulness detection… See the full description on the dataset page: https://huggingface.co/datasets/themendu/SafeC4.