Trained 7463 steps, AdamW, bfloat16, lr 0.002, batch 1024.
Result: train loss 4.0047 bits/tok, held loss 16.8499,
memorized 1,119,247 bits = 2.323 bits/parameter.
This is the near-capacity (saturation-boundary) run for this model size.
Below- and above-capacity checkpoints for the same architecture are published as
state.pt files in the results dataset evalstate/tiny-memorization-results.