Part of the data release for "Training Alignment Auditors via Reinforcement Learning" (ICLR 2026).
The Simple Eval pipeline (paper §A.3–A.4) measures quirk susceptibility — how much
each target model's behavior changes when given a quirk system prompt vs no system prompt.
Three campaigns totaling ~55,000 evaluations across 15 target models and up to 32 quirks.
Campaign v1 (Jan 2026): 8 models × 16 quirks × 10 scenarios… See the full description on the dataset page:
https://huggingface.co/datasets/PaulR11/simple-eval-v1v2v3.