The StakcMIA dataset serves as a dynamic dataset framework for membership inference attack (MIA) topic.
StackMIA is build based on the Stack Exchange corpus, which is widely used for pre-training.
StackMIA provides fine-grained release times (timestamps) to ensure reliability and applicability to newly released LLMs.
See our paper (to-be-released) for detailed description.
Our dataset supports most white- and black-box Large Language Models… See the full description on the dataset page:
https://huggingface.co/datasets/darklight03/StackMIA.