This dataset repository contains a custom 32,768-token BPE tokenizer trained for Small Language Models (SLMs).
It was created using the regex pre-tokenizer rules from Qwen/Qwen2.5-Coder-0.5B.
Total Vocabulary Size: 32,768
BPE Tokens (pure text/bytes): 32,700
Registered Special Tokens: 68 (0 Qwen-inherited + 68 newly injected)
Functional Pad Token:
(ID: 32703)
Alignment Filler Tokens (null pads): 0… See the full description on the dataset page: https://huggingface.co/datasets/rijuludar/slm-tokenizer-32k.