This dataset provides benchmark–target evaluation splits for studying Rubric-Induced Preference Drift (RIPD) in LLM-based evaluation pipelines. The full implementation, rubric search pipeline, and downstream alignment experiments are available in the official GitHub repository:
https://github.com/ZDCSlab/Rubrics-as-an-Attack-Surface.We construct four benchmark–target settings from five widely used… See the full description on the dataset page:
https://huggingface.co/datasets/ZDCSlab/ripd-dataset.