lora_sftmistralai/Mistral-7B-Instruct-v0.3lasgroup/SDPO (slug: sdpo_tooluse)438b979a30de6dfbf3b5a1052e42d8c0453b214d3fp1_sft_math_toolusek5gut2ilstep-00001step-00002step-00004step-00008step-00016step-00032step-00065step-00132revision=... in
AutoModelForCausalLM.from_pretrained / PeftModel.from_pretrained.checkpointing, dataset, evaluation, final_adapter_path, lora, model, optimization, prompt_style, runtime, sdpo, sequence, total_steps