The BookMIA datasets serve as a benchmark designed to evaluate membership inference attack (MIA) methods, specifically in detecting pretraining data from OpenAI models that are released before 2023 (such as text-davinci-003).
The dataset contains non-member and member data:
non-member data consists of text excerpts from books first published in 2023
member data includes text excerpts from older books, as categorized by Chang et al. in 2023.
šā¦ See the full description on the dataset page: https://huggingface.co/datasets/swj0419/BookMIA.