Views
No views yet
fp16 and bf16| Component | Value |
|---|---|
| Layers | 6 |
| Attention Heads | 8 |
| KV Heads | 4 |
| Embedding Dim | 256 |
| Context Length | 512 |
| Vocab Size | 50,304 |
| Attention Type | Multi-Query Attention |
| Norm Type | RMSNorm |
| Position Encoding | Rotary Position Embeddings (RoPE) |
| FFN Activation | SwiGLU (silu) |
safetensors checkpoints → Faster & safer loading.| Property | Value |
|---|---|
| Dataset | FineWeb-MINI |
| Tokens Trained | ~100M |
| Optimizer | AdamW |
| Learning Rate | 6e-4 (cosine decay) |
| Warmup Steps | 100 |
| Batch Size | 64 × 2 grad accum |
| Effective Batch Size | 128 |
| Mixed Precision | fp16 / bf16 (auto-detect) |
| Distributed Training | DDP |
| Logging | Weights & Biases (wandb) |
| Checkpoint Format | .safetensors |
| Step | Filename | Format |
|---|---|---|
| Final | mobile_llm_final.safetensors | safetensors |
| Intermediate | checkpoints/mobile_llm_step_<step>.safetensors | safetensors |