Views
No views yet
trainsft_tulu_tokenize_and_truncate_v1, sft_tulu_filter_v18cd5377c34| Field | Value |
|---|---|
| Base model | hamishivi/Qwen3-8B |
| Local model path | /gpfs/scrubbed/osey/tmax/models/hamishivi/Qwen3-8B |
| Max seq length | 32768 |
| Sequence-parallel size | 1 |
| Per-device batch size | 2 |
| Gradient accumulation | 4 |
| Effective batch size | 128 |
| Packing | true |
| Learning rate | 1e-5 |
| LR scheduler | linear |
| Warmup ratio | 0.1 |
| Epochs | 2 |
| Checkpoint every | 250 steps |
| Seed | 123 |
| Optimizer kernel | Liger |
| Mixed precision | bf16 |
| Gradient checkpointing | true |
| DeepSpeed config | stage2_no_offloading_accelerate.conf |
28 (total 16)117347 on g[007,021]20260516_120450tmax-full-sft (run name: sft_qwen3_8b_tmax_full_0513_2node)sbatch scripts/slurm/sft/qwen/sft_qwen3_8b_tmax_full_0513_2node.sh