The dataset is a relabel dataset of the 22k randomly sampled hhrlhf dataset (11k from helpful-base, 11k from harmless-base).
The annotators are Reward Model trained on this 22k original dataset based on the llama-2-7b .
The annotation python script is as follows:
Unless you fully understand the significance of this dataset, I do not recommend using it lightly.
The re-annotation of the dataset entirely depends on a llama-2-7b reward model, which I have not uploaded yet.
from transformers… See the full description on the dataset page:
https://huggingface.co/datasets/chadlzx/hhrlhf-relabel.