Views
No views yet
| Parameters | 8.4M |
| Layers | 6 |
| d_model | 256 |
| Attention heads | 8 |
| Context length | 512 |
| Vocabulary | 8,192 (BPE ByteLevel) |
| Positional encoding | RoPE |
| Normalization | RMSNorm |
| Activation | SwiGLU |
| Dataset | FineWeb-Edu sample-10BT (~5M tokens) |
| Steps | 1,800 |
| Optimizer | AdamW, cosine LR + warmup |
| Val loss | 5.2764 |
| Perplexity | 195.7 |
| Hardware | Apple Silicon MPS |
1from transformers import PreTrainedTokenizerFast
2tokenizer = PreTrainedTokenizerFast.from_pretrained("REPO_ID")
3print(tokenizer("The study of mathematics").tokens())