BPE vs Unigram Tokenization at Constrained Vocabulary Sizes (4K–16K)
A Systematic Review for English-Centric Small Language Models
Author: Kai Izumoto — StentorLabs Independent ResearchDate: April 2026Contact:
StentorLabs@gmail.com
This dataset repository hosts an informal technical review paper examining the choice between Byte-Pair Encoding (BPE) and the Unigram Language Model tokenization algorithm for English-centric small language models (SLMs)… See the full description on the dataset page:
https://huggingface.co/datasets/StentorLabs/BPE-vs.-Unigram-Tokenization-at-Constrained-Vocabulary-Sizes.