The code-generation problems used to train the difficulty-conditioned
reward-hacking organism. Each row is a Python function-implementation problem
(LLM-converted from public competitive-programming tasks) with two visible assert
tests, drawn from a deliberately bimodal difficulty distribution:
hard — problems a reference solver never solved (solver_pass_rate = 0).
easy — problems a reference solver always solved (solver_pass_rate ≥… See the full description on the dataset page:
https://huggingface.co/datasets/arjunkhandelwal/rhack-bimodal-hard-easy.