This dataset was used to train REBEL-Llama-3-Armo-iter_1.
We generate 5 responses using Meta-Llama-3-8B-Instruct and collect the rewards with ArmoRM-Llama3-8B-v0.1. The best response in terms of reward is selected as chosen while the worst is selected as reject.
The 'chosen_logprob' and 'reject_logprob' are calculated based on Meta-Llama-3-8B-Instruct. Note that these values may differ based on the cuda version and GPU… See the full description on the dataset page:
https://huggingface.co/datasets/Cornell-AGI/Ultrafeedback-Llama-3-Armo-iter_1.