
Sophira-tokenizer-64k-v0 is the first frozen tokenizer artifact for the Sophira project. It is intended for Italian decoder-only language model pretraining.SophiraItalianFM-tokenizer64000huggingface-tokenizersCorpus v0.1 (80% CulturaX Italian subset, 20% PleIAs/Italian-PD)tokenizer-production-1b.manifest.jsonTrue['<bos>', '<eos>', '<unk>', '<pad>']uonlp/CulturaX Italian subset and PleIAs/Italian-PDCorpus v0.180/20 (CulturaX / Italian-PD)1,000,000,000 normalized charactersuonlp/CulturaX Italian subsetPleIAs/Italian-PDsapienzanlp/Minerva-3B-base-v1.0Musixmatch/umberto-commoncrawl-cased-v1gsarti/it5-efficient-small-el32sapienzanlp/modello-italia-9bSophira-tokenizer-64k-v0 achieved the best compression among the successfully evaluated baselines on the held-out Italian slices:
0.1953061.264066| Tokenizer | Mean Tokens/Char | Mean Tokens/Word | All Round-Trip Equal |
|---|---|---|---|
| Sophira-64k | 0.195306 | 1.264066 | no |
| it5-efficient-small-el32 | 0.215350 | 1.393791 | no |
| UmBERTo-commoncrawl-cased-v1 | 0.216695 | 1.402497 | no |
| modello-italia-9b | 0.224812 | 1.455031 | no |
| Minerva-3B-base-v1.0 | 0.240413 | 1.556008 | no |
tokenizer_metadata.json records the frozen artifact metadata.tokenizer_training_config.json records the training configuration.tokenizer_release_manifest.json records checksums and artifact references.tokenizer_evaluation_report.mdapache-2.0.1@misc{peikos2026sophiratokenizer64kv0,
2 title = {Sophira-tokenizer-64k-v0},
3 author = {Peikos, Georgios},
4 year = {2026},
5 publisher = {Hugging Face},
6 howpublished = {\url{https://huggingface.co/Gpeik/Sophira-tokenizer-64k-v0}},
7}