Views
No views yet
meta-llama/Llama-3.2-1B-InstructDPOTrainer with the LLM Judge preference dataset (50 pairs)| Parameter | Value |
|---|---|
| Learning Rate | 5e-5 |
| Batch Size | 4 |
| Epochs | 3 |
| Beta (DPO regularizer) | 0.1 |
| Max Input Length | 1024 tokens |
| Max Prompt Length | 512 tokens |
| Padding Token | eos_token |
llm_judge_preferences.csvprompt, chosen, and rejected columnsLikhith003/dpo-llmjudge-lora-adapter