Views
No views yet
| Metric / Parameter | Multi-Scale v8 | Single-Scale v9 | Matrix SSM v13 (New SOTA) |
|---|---|---|---|
| Architecture | Vector streams + stride-4 Coarse | Single-Scale Vector SSM | Matrix SSM (Quadratic associative state) |
| Parameters | 27,544,336 | 28,662,144 | 28,339,320 |
| Layers | 8 | 16 | 20 (+25% depth) |
| State Capacity | 4,608 floats / layer | 4,608 floats / layer | 6,144 floats / layer (+50% capacity) |
| Vocab Size | 16,384 | 16,384 | 10,240 (parameter-reclaiming) |
| LR Schedule | WSD (Cosine tail) | WSD (Cosine tail) | Pure Cosine Decay |
| QK-Norm | No | No | Yes (Stable projections) |
| Val Loss (Shuffled) | 3.4352 | 3.3719 | 3.2844 |
| Val Perplexity | 31.04 | 29.13 | 26.69 (-8.37% drop vs v9) |
1from transformers import AutoModelForCausalLM, AutoTokenizer
2
3tokenizer = AutoTokenizer.from_pretrained("ecreeth/matrix-ssm-28m-babylm")
4model = AutoModelForCausalLM.from_pretrained(
5 "ecreeth/matrix-ssm-28m-babylm",
6 trust_remote_code=True,
7)
8
9inputs = tokenizer("The cat sat on the", return_tensors="pt")
10outputs = model.generate(**inputs, max_new_tokens=50)
11print(tokenizer.decode(outputs[0]))| Task | Score (Vanilla Baseline) | Score (Multi-Scale v6) | Score (Single-Scale v9) | Score (Matrix SSM v13) |
|---|---|---|---|---|
| BLiMP | 62.87 | 66.78 | 69.99 | 69.72 |
| EWOK (supplement) | 49.64 | 54.33 | 51.54 | 54.46 |
| VQA (EWoK) | 52.76 | 52.35 | 53.33 | 52.86 |
| Entity Tracking | 17.90 | 17.41 | 17.45 | 39.75 |
| Comps | 52.40 | 52.45 | 53.84 | 54.67 |
| Reading (eye tracking) | 0.93 | 1.52 | 0.41 | 0.21 |
| Reading (self-paced) | 0.14 | 0.02 | 0.53 | 0.00 |
| Task | Metric | Score (Vanilla Baseline) | Score (Multi-Scale v6) | Score (Single-Scale v9) | Score (Matrix SSM v13) |
|---|---|---|---|---|---|
| BOOLQ | accuracy | 63.8 | 64.59 | 64.46 | 64.22 |
| MULTIRC | accuracy | 58.5 | 56.93 | Pending / TBD | 57.10 |
| RTE | accuracy | 61.2 | 59.71 | Pending / TBD | 51.08 |
| WSC | accuracy | 63.5 | 63.46 | Pending / TBD | 57.69 |
| MRPC | f1 | 69.6 | 82.25 | Pending / TBD | 81.55 |
| QQP | f1 | 69.6 | 54.98 | Pending / TBD | 55.71 |
| MNLI | accuracy | 43.6 | 44.68 | Pending / TBD | 45.29 (WIN) |
model.step().| Parameter | Value |
|---|---|
| Optimizer | MuonAdamW (Muon for 2D weights, AdamW for embeddings/biases) |
| LR schedule | Cosine Decay (warmup 500 steps, Cosine cooldown) |
| Epochs | 10 epochs (max allowed is 10) |
| Peak LR | 0.0100 (muon), 0.0005 (adamw) |
| Weight decay | 0.1 |
| Batch size | 512 × 1 accum = 512 effective × 512 tokens (262k tokens/step) |
| Total steps | 7,247 |
| GPU | NVIDIA A100 GPU (~3.7h training run) |
| Val loss | 3.2844 (final validation loss on globally shuffled splits) |
| Data cleaning | CHILDES speaker tags, bracket annotations, Wikipedia headers, subtitle formatting, HTML tags filtered |