transformers library:1from transformers import AutoTokenizer
2
3# Load EspBPE-49K directly from Hugging Face Hub
4tokenizer = AutoTokenizer.from_pretrained("DinoResearch/EspBPE-49K")
5
6# Sample text
7text = "Estamos entrenando un tokenizer en español super rápido y eficiente con 49,152 de vocabulario! 🚀"
8
9# Encode text into token IDs
10input_ids = tokenizer.encode(text)
11tokens = tokenizer.convert_ids_to_tokens(input_ids)
12
13print(f"Total Tokens: {len(input_ids)}")
14print("Tokens:", tokens)
15print("Decoded:", tokenizer.decode(input_ids))
16| Text Type | Character Count | Token Count | Compression Ratio |
|---|---|---|---|
| Sports Commentary | 142 chars | 28 tokens | ~5.07 chars / token |
| Encyclopedic History (Wikipedia) | 245 chars | 49 tokens | ~5.00 chars / token |
| General Spanish Sentence | 107 chars | 22 tokens | ~4.86 chars / token |
spa_Latn): Modern, high-quality web crawl text.20231101.es): Encyclopedic knowledge baseline.es): Filtered, clean multilingual web data (OSCAR + mC4).pysentimiento): Informal dialogue, regional accents, and modern social media slang.inconstitucionalidad, extraordinariamente, descentralización, afortunadamente, and Rehabilitación map directly to single Token IDs.á, é, í, ó, ú, ñ) without producing <unk> errors.🤪, 😌, 🇵🇷, 🟢) are learned as native single-token lookup entries.javascript, tensor, Unix, coaxial) alongside regional slang across Spanish-speaking countries.49k)Ġ (ByteLevel space representation)o200k_base tokenizer (used in GPT-4o) across 500 unfiltered Spanish web documents (2,292,109 raw characters) from the FineWeb-2 (spa_Latn) dataset stream.| Metric | 🦖 EspBPE-49K | 🌐 o200k_base (GPT-4o) | Advantage / Delta |
|---|---|---|---|
| Vocabulary Size | 49,152 | 200,019 | ~4x Smaller Table |
| Total Encoded Tokens | 484,581 | 520,509 | -35,928 Tokens (-6.90%) |
| Compression Ratio | 4.730 chars/token | 4.404 chars/token | +0.326 chars/token |
| Document Match Wins | 441 / 500 (88.2%) 🥇 | 50 / 500 (10.0%) | Landslide Victory |
| Draws / Ties | 9 / 500 (1.8%) | 9 / 500 (1.8%) | — |
o200k_base's vocabulary size, saving tens of millions of embedding parameters during pretraining.