Views
No views yet
Qwen/Qwen3-0.6B-Base
(vocab 151,936).model.py and config.json. 1 shared + 16
routed experts top-2, GQA(10/2)×16 layers, partial RoPE, QK-Norm,
SwiGLU, tied embeddings.| key | value |
|---|---|
| param_count_total | 616,471,632 |
| param_count_active | ~203,591,000 (sum(p.numel()) minus inactive routed experts) |
| dataset | openbmb/Ultra-FineWeb (split=en, content) |
| tokenizer | Qwen/Qwen3-0.6B-Base (vocab=151,936) |
| step | 365000 |
| tokens_consumed | 11,960,320,000 |
EVAL_*.md).