Views
No views yet
<unk> fallback, which matters since the training corpus mixes Devanagari (Hindi) and Latin (Portuguese/Spanish) scripts.<|begin_of_text|>, <|end_of_text|>, <|pad|>, <|start_header_id|>/<|end_header_id|>/<|eot_id|>), so the tokenizer is chat-fine-tuning-ready without resizing the embedding table later.pt-fineweb2, es-fineweb2, hi-fineweb2, and hi-sangraha (ai4bharat/sangraha's "verified" split), deduplicated per-language with datatrove (hi's two sources deduplicated against each other, not just internally).1from transformers import AutoTokenizer
2
3tok = AutoTokenizer.from_pretrained("andre15silva/pt-es-hi-tokenizer")
4tok.encode("Olá, ¿cómo estás? नमस्ते")| lang | fertility | compression |
|---|---|---|
| pt | 1.459 | 4.332 |
| es | 1.366 | 4.552 |
| hi | 3.247 | 3.951 |
| overall | 1.773 | 4.264 |