AlphaNeural
llama-3-8b-base-new-dpo-ultrafeedback-4xh200-batch-128-q_t-0.5-s_star-0.4-20260429-032138-margin – Dataset by jackf857 | AlphaNeural AI