Views
No views yet
slm-125m-base (v1) → extended → e2 (2 epochs on the rebuilt 2.5B corpus)
→ e4 (this model — 2 further epochs on the same corpus, 4 epochs total).Ace-2504/slm-125m-e2 for 2 more epochs (9,440 steps, ~4.95B
tokens; ~2.47B per epoch) on the same rebuilt 2.5B-token corpus (stricter OCR
gate, exhausted case-law source). Fresh cosine schedule, peak LR 3e-4 → floor 3e-5,
one A100-40GB. It exists to answer a research question: once continued
pretraining has already recovered, do further epochs on the same in-distribution
data keep improving a small model, or start to overfit and forget?| Domain | loss at e2 start | loss at e4 end | change |
|---|---|---|---|
| SEC filings | 1.7470 | 1.6977 | −2.82% |
| US case law | 2.2828 | 2.2324 | −2.21% |
| FineWeb-Edu | 3.1580 | 3.0987 | −1.88% |
extended run showed on SEC text.slm-125m-base, so token ids mean exactly what they did in v1 training.Ace-2504/slm-125m-base for the
original, or Ace-2504/slm-125m-e2 for the 2-epoch checkpoint.