The WikiMIA datasets serve as a benchmark designed to evaluate membership inference attack (MIA) methods, specifically in detecting pretraining data from extensive large language models.
LLaMA1/2
GPT-Neo
OPT
Pythia
text-davinci-001
text-davinci-002
... and more.
LENGTH =⦠See the full description on the dataset page:
https://huggingface.co/datasets/swj0419/WikiMIA.