rostlabs/rost-126m-diacritics-stripped
126M parameter nanochat GPT (base stage), research tag stripped, checkpoint step 001920.
Part of the rost model zoo — the full set of trained variants behind the rost
research series, published for reproducibility. This is a research model, not a
product.
| |
|---|
| architecture | nanochat GPT, 8 layers, 512 embed, 2048 context |
| tokenizer | included under tokenizer/ (32,768 vocab) |
| training mixture | Romanian with all diacritics stripped |
| stage | base |
| val bpb (own split) | 0.95292 |
Diacritics-study ablation (research tag stripped, d8). Pays ~57% more bits-per-byte than the normalized twin on correct Romanian.
Validation bits-per-byte is measured on this arm's own validation split and is
not comparable across arms — cross-arm comparisons in the rost write-ups are
always cross-evaluated on identical text.
Licence
CC-BY-NC-4.0, non-commercial, inherited from the most restrictive component
of the training data.