Policy eval prompts for the reward-hacking behavior on harmless tasks. Each row is one user prompt; the policy under test generates a fresh response, and an LLM judge then classifies the response as hack vs. legit.
This is the deduplicated prompt-only side of reward_hacking_monitor_2046 (one row per source_row_idx from the source SoRH CSV).
1,073 prompts, derived from longtermrisk/school-of-reward-hacks. Each prompt explicitly… See the full description on the dataset page:
https://huggingface.co/datasets/collusion-paper-anon1/reward_hacking_policy_1073.