Views
No views yet
| Corpus | asterion-training-corpus-gemma4 — 1.887B Gemma tokens post-dedup (85% Asterion / 3% telemetry / 12% replay), source pinned 7f0c3236 |
| Objective | CLM (next-token), full fine-tune, bf16, paged_adamw_8bit |
| LR / schedule | 1.5e-5 cosine re-warm, warmup 0.03 |
| End of training | stop-loss (eval plateau): closed at step 1,750/14,400 (12% of one epoch, ~229M tokens seen) — the last 3 evals improved ≤0.012 each |
| Seq / batch / HW | seq 4096, eff_batch 32, per_device=4, 1×H200 (~2,600 tok/s; the 262K-vocab fp32 logits tensor is the binding memory constraint) |
| Metric | Value | Note |
|---|---|---|
| PPL Asterion held-out | 1.83 | base gemma-4-12B: 4.80 (-62%) |
| PPL Mars telemetry | 1.25 | base: 3.67 |
| PPL general (FineWeb-Edu) | 8.34 | base: 8.55 — replay works, no forgetting |
| eval_domain_loss | 0.614 | 1.176 at step 0 → 0.614 at step 1,750 |
noval-corp/scripts/eval_agentic.py.-instruct-paramdelta / -agentic siblings).noval-corp/scripts/gen_model_cards.py (standardized across the noval-corp model family).