Views
No views yet
| Architecture | Decoder-only transformer (GPT-2 style, pre-LN, tied embeddings) |
| Parameters | 30.3M total (10.6M non-embedding) |
| Layers / heads / width | 6 / 6 / 384 |
| Context length | 1024 |
| Vocabulary | GPT-2 BPE, 50,304 (padded) |
| Precision (training) | fp16 + loss scaling (Tesla T4) |
| Format | safetensors |
| Data | TinyStories, ~400M tokens (GPT-2 tokenizer, uint16 memmap shards) |
| Steps | 6,104 (65,536 tokens/step: batch 8 × 1024 ctx × grad-accum 8) |
| Optimizer | AdamW (β=0.9/0.95, wd 0.1), lr 6e-4, cosine to 10%, 300 warmup steps |
| Hardware | 1× NVIDIA T4 (free Colab), ~2 hours, ~49k tokens/sec |
| Checkpointing | Atomic checkpoints pushed to this repo every 300 steps |
transformers.AutoModel. Weights are standard safetensors
(see config.json in the checkpoint folder for the architecture: 6 layers,
6 heads, width 384, GPT-2 BPE tokenizer). The training and inference code
is not yet public; it will be released alongside a later model in the
series.Once upon a time there was a little robot. He was very happy and liked to roll with his friends. But one day, he rolled too fast and fell into a big puddle. He tried to roll out of his wet puddle, but he couldn't. He was stuck and couldn't get out. Luckily, a kind little girl saw the robot and knew just what to do. [...] From then on, the robot was extra careful.