This model is a fine-tuned version of
meta-llama/Meta-Llama-3.1-8B-Instruct on the prm_conversations_math-l5-train-full-ref-mix_gpt-4o-mini_rm_best_of_16_with_ref_with_hint_all_full_solution_as_ref dataset.
It achieves the following results on the evaluation set: