AlphaNeural
Llama-3-8B-Instruct-SPPO-Iter3-gp-8b-gpm-reg0.5-sppo-reversekl-table – AI Model by RegularizedSelfPlay | AlphaNeural AI