AlphaNeural
dpo_helfulhelpful_gamma0.0_beta0.1_subset20000_modelmistral7b_maxsteps5000_bz8_lr1e-05 – AI Model by Holarissun | AlphaNeural AI