Trained 7330 steps, AdamW, bfloat16, lr 0.002, batch 512.
Result: train loss 3.7714 bits/tok, held loss 18.0931,
memorized 4,626,295 bits = 2.473 bits/parameter.
This is the near-capacity (saturation-boundary) run for this model size.
Below- and above-capacity checkpoints for the same architecture are published as
state.pt files in the results dataset evalstate/tiny-memorization-results.