This dataset consists of the attack samples used for the paper "How Much Do Code Language Models Remember? An Investigation on Data Extraction Attacks before and after Fine-tuning"
We have two splits:
The fine-tuning attack, which consists of selected samples coming from the fine-tuning set
The pre-training attack, which consists of selected samples coming from the TheStack-v2 on the Java section
We have different splits depending on the duplication rate of the samples: