harshbhatt7585/tinygroot-sft-mtp-heads-20260527-2
tinyGroot checkpoint uploaded from the training pipeline.
Checkpoint
- Step:
974
- Tokenizer:
nanochat
- Layers:
12
- Hidden size:
768
- Heads:
6
- MTP heads:
2
- Max sequence length:
2048
- Source run:
sft-mtp-heads-20260527-2
Files
model.pt: model state_dict (mmap+weights_only=True friendly).
optimizer.pt: optimizer and grad-scaler state (only needed to resume training).
meta.json: step, model config, tokenizer type, and original training args.
tokenizer_hf/tokenizer.json: tokenizer used by this checkpoint.
Load this checkpoint with the tinyGroot codebase, not the Transformers AutoModel API.