AlphaNeural
Llama-3-8B-Instruct-SPPO-Iter2-gp-8b-gpm-reg0.05-sppo-reversekl-table – AI Model by RegularizedSelfPlay | AlphaNeural AI