Views
No views yet
meta-llama/Llama-3.2-1B-InstructDPOTrainer from TRL on the PairRM preference dataset consisting of 500 preference pairs.| Setting | Value |
|---|---|
| Learning Rate | 3e-5 |
| Batch Size | 4 |
| Training Epochs | 3 |
| DPO Beta | 0.1 |
| Max Sequence Length | 1024 tokens |
| Max Prompt Length | 512 tokens |
| Padding Token | EOS Token |
pairrm_preferences.csvprompt, chosen, and rejected.sahithimuppavaram/dpo-pairrm-lora-adapter