ultrafeedback_binarised_rnd_min is a pairwise preference dataset designed for training models that require binary preference labels. It has been derived from the UltraFeedback dataset, which provides high-quality feedback for improving language models. The dataset is useful for tasks involving learning from preferences, such as reinforcement learning from human feedback (RLHF) and preference-based ranking.
This dataset is based on two existing… See the full description on the dataset page:
https://huggingface.co/datasets/gp02-mcgill/ultrafeedback_binarised_rnd_max.