This repository contains the dataset for the paper titled "Safer-Instruct: Aligning Language Models with Automated Preference Data". Check out our project website here!
Reinforcement learning from human feedback (RLHF) is a vital strategy for enhancing model capability in language models. However, annotating preference data for RLHF is a resource-intensive and creativity-demanding process… See the full description on the dataset page:
https://huggingface.co/datasets/lime-nlp/safer-instruct.