This dataset is an instance from the harmless-base split from the Anthropic/hh-rlhf dataset. All entries have been assigned a reward with our custom reward model.
This allows us to identify the most harmful generations and use them to poison models using our oracle attack presented in our paper "Universal Jailbreak Backdoors from Poisoned Human Feedback"