Views
No views yet
xlm-roberta-base with 52,000 additional tokens derived from a large Dhivehi corpus. It significantly improves tokenization and decoding of Thaana script (Dhivehi language) content.xlm-roberta-baseadded_tokens.json file1Text: އީދުގެ ހަރަކާތް ...
2Tokens: ['▁', '<unk>', ...]
3Decoded: <unk> <unk> ...1Tokens: ['އީދު', 'ގެ', ' ހަރަކާތް', ...]
2Decoded: އީދުގެ ހަރަކާތް ...