This model is a fine-tuned version of
/home/shutingw/.cache/huggingface/hub/models--meta-llama--Meta-Llama-3-8B-Instruct/snapshots/5f0b02c75b57c5855da9ae460ce51323ea669d8a on the /data/user_data/shutingw/wentaos/Optima/my_datasets/trival_qa_sft_dpo_DI/sft/iteration_0 dataset.
It achieves the following results on the evaluation set: