Views
No views yet
kowo-co/babble (pretrain_hf.py + the default 34.1M config).| Parameters | 34,096,128 |
| Layers / width / heads | 8 / 512 / 8 |
| Context | 1024 tokens |
| Tokenizer | byte-level BPE, 16,384 tokens |
| Data | openbmb/Ultra-FineWeb-L1 |
| Train split | CC-MAIN-2025-51 |
| Val split | CC-MAIN-2025-47 (disjoint crawl) |
| Tokens trained | 600,206,202 |
| Hardware | 1× NVIDIA H200 (Hugging Face Jobs) |
| Wall clock | ~51 minutes end-to-end (~46 min of training at ~217k tok/s) |
| Final train loss | 2.531 |
| Final val loss | 2.472 |
latest.pt — checkpoint (model state dict, config, optimizer, step/token counts)tokenizer.json — BPE merge list, loadable with babble.subword.BPETokenizer.from_jsonloss.jsonl — per-checkpoint train/val loss, throughput, and samplesthe cat → the catastrophic system and its adaptability to manage the catastrophic system is the case with a significant surgeIn the beginning → In the beginning of the New Jersey Law and Law, “law enforcement of the Law and Law and Law in the lawScientists have discovered → Scientists have discovered that the current market is expected to take a long way to see how we look at the new marketThe weather today is → The weather today is about to create a more sustainable and sustainable future. It’s a way to make the most of your time andpretrain_hf.py (bf16 AMP, AdamW, cosine LR after warmup). Job: ProCreations/6a893d3e7c5c7dd37923450f.