Views
No views yet
| Parameter | Value |
|---|---|
| Learning Rate | 2e-5 |
| Weight Decay | 1e-4 |
| Epochs | 2 |
| Global Batch Size | 64 (per_device=1, grad_accum=2) |
| Micro-batch Size per GPU | 1 |
| Max Sequence Length | 262144 tokens |
| Optimizer | AdamW (β₁=0.9, β₂=0.95) |
| LR Scheduler | Cosine with 10% warmup |
| Gradient Clipping | 1.0 |
| Packing | Enabled |
| Precision | BF16 |
| Item | Value |
|---|---|
| GPU | H200 |
| Nodes | 8 |
| GPUs per Node | 8 |
| Total GPUs | 64 |
1from transformers import AutoModelForCausalLM, AutoTokenizer
2
3model = AutoModelForCausalLM.from_pretrained(ZhuofengLi/qwen3.5-9b-nemotron-sft-ckpt200, torch_dtype=auto)
4tokenizer = AutoTokenizer.from_pretrained(ZhuofengLi/qwen3.5-9b-nemotron-sft-ckpt200)