Trained 15698 steps, AdamW, bfloat16, lr 0.002, batch 512.
Result: train loss 5.8075 bits/tok, held loss 16.6023,
memorized 14,622,163 bits = 1.847 bits/parameter.
This is the near-capacity (saturation-boundary) run for this model size.
Below- and above-capacity checkpoints for the same architecture are published as
state.pt files in the results dataset evalstate/tiny-memorization-results.