This repository contains data from 22 different domains of the PILE, divided into train and val sets. The data is in the form of a JSON file, with each entry containing the raw text, as well as various kinds of perturbations applied to it. The dataset is used to facilitate privacy research in language models, where the perturbed data can be used as reference detect the presence of a particular dataset in the training data of a… See the full description on the dataset page:
https://huggingface.co/datasets/bxiong/dataset_inference.