This repository contains a preference dataset designed for Direct Preference Optimization (DPO) training. The preferences were generated programmatically using the llm-blender/PairRM reward model.
This dataset was created as part of the "Preference Dataset Collection and DPO Training" project to fine-tune meta-llama/Llama-3.2-1B-Instruct.
Fine-tuned Model: NilayR/llama32-dpo-pairrm
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/NilayR/pairrm-preferences-llama32.