Training data for ACL 2025 paper "Aligning Large Language Models with Implicit Preferences from User-Generated Content" (
https://arxiv.org/abs/2506.04463)
The file "Mistral_single-K5-prometheus-reward_ref_run8.json" file is directly used for model training, DPO by default.
The file "mixv2-gen_filter_data.jsonl" file is the raw file that contains user query we synthesized and filter from the user-generated content.
If you find this resource useful, please kindly cite our paper:… See the full description on the dataset page:
https://huggingface.co/datasets/Zhaoxuan/PUGC_data.