The pretrained base model from a nanochat-style run with
ohara. 124M effective parameters, trained
from scratch on 1.48B tokens of ClimbMix on 2xA100-80GB in about 3.3 hours.
This is a base language model, not a chat model. It continues text; it does
not answer questions and it never stops on its own. Given "Why is the sky blue?"
it will happily continue into an unrelated encyclopedia article. If you want
something that responds to you, use
ohara-chat-d12, which is this
checkpoint after supervised finetuning.
Use this one if you want to run your own finetuning from a base.
Architecture
Llama-style decoder: RoPE, SwiGLU, RMSNorm, untied embeddings, nanochat-style
initialization, trained with Muon (matrices) + AdamW (embeddings and scalars).
Effective / total params
123.5M / 162.2M
Layers / hidden / heads
12 / 768 / 6
Context length
2048
Vocabulary
50,265 (gpt-neo-125m + 8 reserved chat tokens)
The 8 conversation special tokens are reserved in the vocabulary but never seen
during pretraining, so a finetune can use them without resizing the embedding.
Training
2,827 steps over 1.48B tokens, batch 524,288 tokens, Muon at lr 0.02 under a
warmup-stable-decay schedule, bf16 mixed precision.
step
val bits/byte
750
1.0625
1250
1.0057
2000
0.9458
2827
0.9062
Final val loss 2.9153, next-token accuracy 43.6%.
For scale: nanochat reaches GPT-2 grade at 0.718 bpb using roughly 28x more
compute than the 1.4e18 FLOPs spent here. This is an honestly small model.