Views
No views yet
meta-llama/Llama-3.2-1B-InstructDPOTrainer using the LLM Judge preference dataset with 50 human-labeled comparisons.| Setting | Value |
|---|---|
| Learning Rate | 5e-5 |
| Batch Size | 4 |
| Training Epochs | 3 |
| DPO Beta | 0.1 |
| Max Sequence Length | 1024 tokens |
| Max Prompt Length | 512 tokens |
| Padding Token | EOS Token |
llm_judge_preferences.csvprompt, chosen, and rejected labels.sahithimuppavaram/dpo-llmjudge-lora-adapter