Per-round trajectory and held-out samples from an iterative self-training loop on GSM8K (Qwen2.5-7B-Instruct, LoRA DPO, 6 rounds). Same loop every arm; only the preference-label source differs. This dataset is the self arm: pairs labelled by the model judging its OWN answers (pairwise, both-orders position-bias filter).
Study question: does a model training on its own judgment collapse? Short answer on verifiable math: the reward can be hacked, but… See the full description on the dataset page:
https://huggingface.co/datasets/yavuz-ai/self-reward-collapse-self.