Views
No views yet
| Parameter | Value |
|---|---|
| Architecture | GPT-2 small (124M params) |
| Layers | 12 |
| Hidden dim | 768 |
| Attention heads | 12 |
| Context length | 512 |
| FFN dim | 2048 |
| Vocab size | 50,257 (GPT-2 BPE) |
| RoPE theta | 10,000 |
| Parameter | Value |
|---|---|
| Hardware | 2x NVIDIA RTX A4000 (16 GB each) |
| Precision | bfloat16 |
| Training steps | 12,000 |
| Effective batch size | 256 sequences/step |
| Tokens per step | 131,072 |
| Total tokens seen | ~1.57B |
| Training time | ~14 hours |
| Best val loss | 2.856 |
| Step | Val Loss |
|---|---|
| 1,000 | 4.032 |
| 2,000 | 3.683 |
| 3,000 | 3.507 |
| 4,000 | 3.379 |
| 5,000 | 3.273 |
| 6,000 | 3.202 |
| 7,000 | 3.101 |
| 8,000 | 3.030 |
| 9,000 | 2.982 |
| 10,000 | 2.927 |
| 11,000 | 2.883 |
| 12,000 | 2.856 |
BasicsTransformerLM) with RoPE embeddings and SwiGLU FFN. Load with:1import torch
2import json
3
4config = json.load(open("model_config.json"))
5state_dict = torch.load("model.pt", map_location="cpu")