This repository hosts a copy of the widely used WikiMIA dataset,a benchmark designed to evaluate membership inference attack (MIA) methods—specifically for detecting whether a piece of text was seen during the pretraining of Large Language Models (LLMs).
WikiMIA is commonly used in data contamination / pretraining data detection research, including the paper “Detecting Pretraining Data from Large Language Models” (arXiv:2310.16789).
Contents… See the full description on the dataset page: https://huggingface.co/datasets/S3IC/wikimia.