This dataset was used to train REBEL-Llama-3-Armo-iter_2.
We generate 5 responses using REBEL-Llama-3-Armo-iter_1 and collect the rewards with ArmoRM-Llama3-8B-v0.1. The best response in terms of reward is selected as chosen while the worst is selected as reject.
The 'chosen_logprob' and 'reject_logprob' are calculated based on REBEL-Llama-3-Armo-iter_1. Note that these values may differ based on the cuda version and GPU… See the full description on the dataset page:
https://huggingface.co/datasets/Cornell-AGI/Ultrafeedback-Llama-3-Armo-iter_2.