Training: ~8B tokens of pre-1900 English text, FP8 (tensorwise), Muon+AdamW optimizer
Final val BPB: 1.211
Checkpoint Contents
model_007226.pt # Model weights (4.9 GB)
meta_007226.json # Training config and metadata
optim_007226_rank*.pt # Optimizer state, 8 FSDP shards (for resuming training)
tokenizer/ # BPE tokenizer (tiktoken format) + token byte counts
nanochat/ # Source code to load and run the model
eval_results.csv # Benchmark eval results at this checkpoint
Quick Start
python
1import torch
2from nanochat.gpt import GPT, GPTConfig
3from nanochat.tokenizer import RustBPETokenizer
45# Load tokenizer6tokenizer = RustBPETokenizer.from_directory("tokenizer")78# Load model9import json
10withopen("meta_007226.json")as f:11 meta = json.load(f)1213config = GPTConfig(**meta["model_config"])1415with torch.device("meta"):16 model = GPT(config)17model.to_empty(device="cuda")18model.init_weights()1920state_dict = torch.load("model_007226.pt", map_location="cuda")21state_dict ={k.removeprefix("_orig_mod."): v for k, v in state_dict.items()}22model.load_state_dict(state_dict, strict=True, assign=True)23model.eval()2425# Generate26bos = tokenizer.get_bos_token_id()27tokens = tokenizer.encode("It was a dark and stormy night", prepend=bos)28with torch.amp.autocast(device_type="cuda", dtype=torch.bfloat16):29for token in model.generate(tokens, max_tokens=100, temperature=0.8):30print(tokenizer.decode([token]), end="", flush=True)
Dependencies
torch>=2.9
tiktoken
rustbpe
Eval Results (step 7226)
Task
Accuracy
Centered
hellaswag
0.318
0.091
arc_easy
0.411
0.215
lambada_openai
0.332
0.332
piqa
0.586
0.172
winograd
0.674
0.348
copa
0.570
0.140
CORE
0.126
Training
Trained with the nanochat framework using 8x H100 GPUs with FSDP.
To resume training, load the optimizer shards (optim_007226_rank*.pt) — one per FSDP rank.