The UltraFeedback Binarized Dataset is a high-quality preference dataset designed for aligning large language models (LLMs) through preference learning and reinforcement learning from human feedback (RLHF).
Each record contains a prompt and two candidate responses — chosen and rejected — produced by assistant models, along with quality scores that indicate human or model-based preferences.
This processed version extracts… See the full description on the dataset page: https://huggingface.co/datasets/liavonpenn/Processed_UltraFeedback_Binarized.