Views
No views yet
| Parameter | Value |
|---|---|
| Base model | Qwen/Qwen2.5-1.5B-Instruct |
| Method | Algorithmic Template SFT |
| Framework | LLaMA-Factory (SFT stage) |
| LoRA rank | 64 |
| LoRA target | all linear layers |
| Learning rate | 1e-4 |
| Epochs | 3 |
| Batch size | 4 (grad accum 4) |
| Cutoff length | 32,768 tokens |
| Training data | 5,000 deterministic carry-propagation traces (d4: 3-digit x 2-3 digit multiply) |
| Split | Accuracy |
|---|---|
| Test (in-distribution) | 92.6% |
| Harder variant | 21.2% |
| Structural OOD | 0.0% (chain operations) |
1from peft import PeftModel
2from transformers import AutoModelForCausalLM, AutoTokenizer
3
4base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")
5model = PeftModel.from_pretrained(base, "reasoning-degeneration-dev/algo-sft-long-arithmetic-standard")
6tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")