Views
No views yet
⚠️ This model is not optimized for performance or real-world use.
| Parameter | Value |
|---|---|
| Layers | 20 |
| Attention Heads | 10 |
| KV Heads | 10 |
| Embedding Size | 1280 |
| Head Dimension | 128 |
| Context Length | 2048 tokens |
| Vocabulary Size | 32,768 |
| Window Pattern | L |
| Parameter | Value |
|---|---|
| Optimizer | Likely AdamW (not explicitly recorded) |
| Embedding LR | 0.3 |
| Unembedding LR | 0.008 |
| Matrix LR | 0.02 |
| Scalar LR | 0.5 |
| Weight Decay | 0.28 |
| Warmup Steps | 100 |
| Warmdown Ratio | 0.85 |
| Final LR Fraction | 0.05 |
| Metric | Value |
|---|---|
| Validation BPB | 0.8325 |
| Best Val BPB | 0.7318 |
| Train Loss (smoothed) | 2.6494 |
Explain recursion simply:Recursion is when something calls itself but it can be confusing and sometimes it keeps going because it is repeating the same idea again and again.from transformers import pipeline
generator = pipeline("text-generation", model="your-username/nanochat-10k")
print(generator("Explain recursion simply:", max_length=50))