Nepali BPE Tokenizer
This tokenizer was trained on the bashyaldhiraj2067/500k_copy_error_dataset, which contains Nepali text data with both correct and incorrect examples.
Details
- Tokenizer type: BPE (Byte-Pair Encoding)
- Vocabulary size: 50,000
- Minimum frequency: 2
- Special tokens:
, , , ,