This dataset contains the final preference pairs used for the Direct Preference Optimization assignment in this repository.
Base instruction source: GAIR/lima
Candidate generator: local Qwen/Qwen2.5-7B-Instruct
Preference ranker: local llm-blender/PairRM
Sample 50 instructions from the local LIMA training split with seed 42.
Generate 5 candidate responses per instruction with Qwen2.5-7B-Instruct.… See the full description on the dataset page:
https://huggingface.co/datasets/ITBill/INFH-6000Q-dpo-preference-dataset.