Views
No views yet
<bos> {user message} <sep> {reply} <eos><bos> your message <sep> and sample until <eos>. Multi-turn: put earlier turns in the prompt (newlines), last user message at the end, then <sep>.| Parameters | 34,096,128 (8 layers, 512 wide, 8 heads, context 1024) |
| Tokenizer | same byte-level BPE as the pretrain (16,384 tokens) |
| Objective | SFT, assistant tokens only |
| Assistant tokens | 150,001,494 |
| Steps | 36,017 |
| Hardware | 1× NVIDIA H200 (Hugging Face Jobs) |
| Wall clock | ~73 minutes (~72 min of training at ~34.6k target-tok/s) |
| Val loss | 3.05 → 2.590 |
hey → Hii hruwhat's up → I just got on matelol → I think they both are too smallcan you help me → what level is thatlatest.pt — checkpoint (model, config, optimizer, stage: posttrain-discord)tokenizer.json — BPE merges (babble.subword.BPETokenizer.from_json)loss.jsonl — val loss, throughput, and samples per checkpointrun_meta.json — dataset / LR / batch used for this run