Views
No views yet
Msok99/km-improved-22k-v4| Feature | Description |
|---|---|
| Extended Vocabulary | 32,000 tokens for higher domain coverage |
| Improved Context Retention | Keeps compound and rare words intact |
| Reduced Fragmentation | Fewer subword splits across long sentences |
| Perfect Decode Fidelity | 100% reversible encoding/decoding |
| Broad Domain Corpus | Includes academic, scientific, literary, and technical texts |
| Category | Avg Tokens | Chars/Token |
|---|---|---|
| Formal News | 13.6 | 4.19 |
| Technology / Scientific | 10.8 | 5.32 |
| Culture / History | 11.0 | 4.58 |
| Education / Academic | 9.4 | 5.44 |
| Mixed Texts | 12.2 | 3.86 |
| Overall Efficiency | — | ≈4.0 chars/token |
18k or 22k models)Msok99/km-improved-22k-v4.Msok99/lfm2-khmer-merged-18k.1from transformers import AutoTokenizer
2
3tokenizer = AutoTokenizer.from_pretrained("Msok99/km-improved-32k")
4
5text = "សង្គ្រាមត្រជាក់មានឥទ្ធិពលដល់នយោបាយពិភពលោក។"
6tokens = tokenizer.tokenize(text)
7print(tokens)
8print(tokenizer.decode(tokenizer.encode(text)))