This is the data mixture used for the reward model weqweasdas/RM-Mistral-7B, trained with the script
https://github.com/WeiXiongUST/RLHF-Reward-Modeling .
Also see a short blog for the training details (data mixture, parameters...):
https://www.notion.so/Reward-Modeling-for-RLHF-abe03f9afdac42b9a5bee746844518d0
If you have any question with this reward model and also any question about reward modeling, feel free to drop me an… See the full description on the dataset page:
https://huggingface.co/datasets/living-box/preference_dataset_mixture2_and_safe_pku.