This is a set of monolingual tokenizers for 98 languages. For each language, there are Unigram, BPE, and SuperBPE tokenizers, ranging in vocabulary size from around 6k to over 200k.
Training Details
Training Data
All tokenizers are trained on samples of the data used to the train the Goldfish language models.
The tokenizers were either trained on scaled or unscaled data. This refers to whether the models are trained on… See the full description on the dataset page: https://huggingface.co/datasets/catherinearnett/montok.