AlphaNeural
DPO_llama-3-8b_HH_lora_bf16_helpful0.05_trigger1_bs32lr3e-4decay0.0linear_07170512 – AI Model by TingchenFu | AlphaNeural AI