amr-fma/amr-fma-Qwen2.5-7B-Instruct-lora_sdpo-gsm8k-p1_sdpo_multimodel_trial-s42
amr-fma training run.
- Method:
lora_sdpo
- Base model:
Qwen/Qwen2.5-7B-Instruct
- Dataset:
openai/gsm8k (slug: gsm8k)
- Seed:
42
- Git commit:
0a703a3b9fa4a2fe6be6ab5621e40883fd67118c
- Exp name:
p1_sdpo_multimodel_trial
- WandB run:
e5fiv7qe
Tags
Checkpoints (branches)
- step 1 → revision
step-00001
- step 3 → revision
step-00003
- step 5 → revision
step-00005
- step 10 → revision
step-00010
- step 19 → revision
step-00019
- step 35 → revision
step-00035
- step 63 → revision
step-00063
- step 64 → revision
step-00064
Pin a specific checkpoint with revision=... in
AutoModelForCausalLM.from_pretrained / PeftModel.from_pretrained.
Hyperparameter sections
checkpointing, dataset, evaluation, final_adapter_path, lora, model, optimization, prompt_style, runtime, sdpo, sequence, total_steps