TokSuite–ByT5 is part of TokSuite, a controlled suite of language models designed to isolate and measure the impact of tokenizer choice on language model behavior.
This model is architecturally identical to the other 13 TokSuite models and differs only in its tokenizer. All TokSuite models share:
the same model architecture,
the same training data,
the same training budget,
and a shared initialization for overlapping vocabulary items.
As a result, any behavioral differences observed between TokSuite models can be attributed directly to tokenizer design, rather than confounding factors such as data scale or optimization.
Tokenizer
Tokenizer: ByT5
Tokenization method: Byte-level
Vocabulary size: 259
Out-of-vocabulary handling: Byte-based (no OOVs)
Language coverage: Language-agnostic
Pretokenization source: None (raw bytes)
Processing details:
Numbers: Represented as raw UTF-8 bytes
Contractions: N/A
Unicode normalization: None
Whitespace / boundary markers: N/A
Zerowidth chars: 3 Bytes
Why ByT5?
ByT5 represents an extreme point in the tokenizer design space: instead of learning subword units, it operates directly on bytes. This guarantees complete coverage of Unicode text and eliminates vocabulary fragmentation issues common in subword tokenizers.
Because all models are trained with a fixed token budget, the byte-level tokenizer processes significantly more raw text (in bytes) than subword-based tokenizers. This reflects a realistic trade-off between compression efficiency and robustness.
Training Procedure
Training steps: 100,000
Sequence length: 4096
Batch size: 256 sequences
Optimizer: AdamW
Peak learning rate: 1e-3
Schedule: Cosine decay with 2,000 warm-up steps
Weight decay: 0.1
Evaluation
Canonical Benchmarks
The model was evaluated on standard base-language-model benchmarks:
HellaSwag
ARC
PIQA
XNLI
TokSuite Logo
These evaluations verify that the model exhibits reasonable base language modeling behavior at its scale and training budget.
TokSuite Robustness Benchmark
TokSuite–ByT5 is evaluated on the TokSuite multilingual robustness benchmark, which probes real-world perturbations including:
orthographic and spelling errors,
diacritics presence/absence,
keyboard and input-method noise,
Unicode styling and homoglyphs,
OCR and spacing artifacts,
LaTeX and STEM formatting.
Tokenization Robustness under Multilingual Text Perturbations
Values represent relative performance drop, computed as (Acc_clean − Acc_perturbed) / Acc_clean, where lower values indicate greater robustness.
Perturbation types include:
Input: non-native keyboard input and romanization
Diacr.: optional diacritics
Orth.& Gram.: orthographic and grammatical errors
Morph: morphological variations including derivations, inflections, and contractions
Noise: homoglyph substitutions, OCR artifacts, typos, and spacing errors
LaTeX: LaTeX-style mathematical formatting
STEM: scientific diagrams and notational conventions
Unic.: Unicode styling characters
NEN denotes non-English inputs and EN denotes English inputs. The Avg column reports the average relative performance drop across all perturbation categories.
Model
Input (NEN)
Diacr. (NEN)
Orth. & Gram. (EN)
Orth. & Gram. (NEN)
Morph (EN)
Morph (NEN)
Noise (EN)
Noise (NEN)
LaTeX (EN)
STEM (EN)
Unic. (EN)
Avg ↓
TokenMonster
0.23
0.33
0.08
0.01
0.23
-0.07
0.10
0.18
0.21
0.10
0.51
0.17
XGLM
0.34
0.49
0.10
0.11
0.25
0.07
0.12
0.22
0.29
0.29
0.11
0.22
BLOOM
0.30
0.34
0.13
0.07
0.18
0.11
0.18
0.18
0.24
0.11
0.57
0.22
ByT5
0.30
0.44
0.04
0.06
0.27
0.04
0.14
0.18
0.17
0.29
0.53
0.22
Comma
0.28
0.43
0.05
0.07
0.18
0.00
0.11
0.20
0.23
0.29
0.61
0.22
mBERT
0.33
0.44
0.11
0.11
0.23
0.06
0.18
0.22
0.14
0.22
0.61
0.24
GPT-4o
0.30
0.51
0.08
0.05
0.21
0.05
0.16
0.19
0.24
0.33
0.55
0.24
GPT-2
0.34
0.46
0.07
0.10
0.25
0.06
0.14
0.21
0.24
0.35
0.53
0.25
Phi-3
0.33
0.46
0.16
0.09
0.27
0.08
0.17
0.21
0.24
0.22
0.55
0.25
Gemma-2
0.32
0.42
0.14
0.15
0.24
0.03
0.16
0.25
0.22
0.36
0.57
0.26
Qwen-3
0.36
0.42
0.14
0.11
0.25
0.06
0.16
0.23
0.26
0.29
0.57
0.26
Llama-3.2
0.33
0.55
0.11
0.10
0.25
0.08
0.15
0.24
0.17
0.30
0.59
0.26
Aya
0.31
0.46
0.14
0.10
0.22
0.03
0.19
0.25
0.21
0.38
0.58
0.26
Tekken
0.33
0.47
0.18
0.03
0.31
0.10
0.14
0.21
0.27
0.43
0.54
0.27
Avg
0.31
0.44
0.11
0.08
0.24
0.04
0.15
0.21
0.22
0.28
0.53
0.24
Tokenizer-Specific Findings
In TokSuite evaluations, the ByT5 tokenizer exhibits:
Strong robustness to:
multilingual noise,
OCR errors,
keyboard perturbations,
Unicode formatting and homoglyphs.
Particularly strong performance in Turkish and Chinese settings, including romanized inputs.
Minimal catastrophic failure modes, as byte corruption results in predictable token sequences rather than invalid subwords.
However, this robustness comes at a cost:
higher sequence lengths,
reduced tokenization efficiency,
increased computational overhead per example.
These findings highlight a fundamental robustness–efficiency trade-off in tokenizer design.
Intended Use
This model is intended for:
research on tokenization and robustness,
multilingual NLP analysis,
controlled ablation studies,
benchmarking tokenizer behavior under noise.
It is not instruction-tuned, aligned, or optimized for deployment.
Limitations
Trained on a limited set of five languages.
Not optimized for instruction following or dialogue.
Fixed token budget constrains exposure to raw text depending on tokenization efficiency.
Intended strictly for research purposes.
Ethical Considerations
TokSuite models are released strictly for research purposes.
They inherit biases present in large-scale web data and should not be used in high-stakes or sensitive applications without additional alignment and evaluation.
Citation
If you use this model, please cite:
bibtex
1@article{toksuite2025,
2 title={TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior},
3 author={Altıntaş, Gul Sena and Ehghaghi, Malikeh and Lester, Brian and Liu, Fengyuan and Zhao, Wanru and Ciccone, Marco and Raffel, Colin},
4 year={2025},
5 arxiv={https://arxiv.org/abs/2512.20757},
6}