lora_sdpoallenai/OLMo-2-1124-7B-Instructopenai/gsm8k (slug: gsm8k)420a703a3b9fa4a2fe6be6ab5621e40883fd67118cp1_sdpo_multimodel_trial32um78i6step-00001step-00003step-00005step-00010step-00019step-00035step-00063step-00064revision=... in
AutoModelForCausalLM.from_pretrained / PeftModel.from_pretrained.checkpointing, dataset, evaluation, final_adapter_path, lora, model, optimization, prompt_style, runtime, sdpo, sequence, total_steps