Views
No views yet
| Component | Details |
|---|---|
| Parameters | 1,185M |
| Dimensions | 2048 |
| Layers | 24 |
| Vocab size | 50304 |
| Attention | GQA — 16 heads, 4 KV heads, 128 head dim, QK-norm |
| FFN | SwiGLU, 2.667x expansion |
| Residual | Standard pre-norm (RMSNorm) |
| Optimizer | AdamW (LR 3e-4), cosine schedule, 500-step warmup |
| Benchmark | Goedel-Baseline-1B (1,185M) | Goedel-mHC-1B (1,009M) |
|---|---|---|
| BPB (wikitext-2) | 1.130 | 1.087 |
| val_loss (FineWeb-Edu) | 2.686 | 2.645 |
| HellaSwag | 36.2% | 39.7% |
| ARC-Easy | 52.8% | 57.8% |
| ARC-Challenge | 23.9% | 24.3% |
| WinoGrande | 53.1% | 54.9% |
1model:
2 dim: 2048
3 n_layers: 24
4 vocab_size: 50304
5
6attention:
7 type: gqa
8 num_heads: 16
9 num_kv_heads: 4
10 head_dim: 128
11 qk_norm: true
12 rope_theta: 10000
13
14ffn:
15 type: swiglu
16 intermediate_mult: 2.667
17
18residual:
19 type: prenorm
20
21optim:
22 type: adamw
23 lr: 3.0e-4
24 scheduler: cosine
25 warmup_steps: 500
26 weight_decay: 0.1
27 max_grad_norm: 1.0
28
29training:
30 tokens: 20_000_000_000
31 batch_size: 8
32 seq_len: 4096
33 grad_accum_steps: 6
34 liger: true
35 compile: true
36 compile_mode: max-autotune-no-cudagraphs
37
38data:
39 shard_dir: data/fineweb_edu