lora_sdpomistralai/Mistral-7B-Instruct-v0.3DigitalLearningGmbH/MATH-lighteval (slug: math)428b979a30de6dfbf3b5a1052e42d8c0453b214d3fp1_sdpo_math_l3plus02olkt1mstep-00001step-00003step-00006step-00012step-00022step-00042step-00079step-00080revision=... in
AutoModelForCausalLM.from_pretrained / PeftModel.from_pretrained.checkpointing, dataset, evaluation, final_adapter_path, lora, model, optimization, prompt_style, runtime, sdpo, sequence, total_steps