Views
No views yet
transformers and tokenizers libraries installed:pip install transformers tokenizers<unk>, <s>, and </s> for managing unknown tokens, and marking the beginning and end of sequences, respectively.1messages = [
2 {"role": "system", "content": "ଆପଣ ଜଣେ ସହାୟକ ଏବଂ ଆପଣଙ୍କର ଭଲ ଐତିହାସିକ ଜ୍ଞାନ ଅଛି।"},
3 {"role": "user", "content": "ତୁମେ କେଉଁଠାରୁ ଆସିଛ?"},
4 {"role": "assistant", "content": "ମୁଁ ପୃଥିବୀରୁ ଆସିଛି"}
5]1Tokenized input (ID list): [414, 329, 225, 316, 221, 318, 309, 274, 531, 222, 303, 596, 1506, 223, 237, 221, 238, 223, 226, 272, 222, 312, 221, 227, 676, 295, 276, 229, 231, 225, 232, 225, 292, 275, 300, 221, 224, 229, 349, 223, 255, 33, 249, 463, 236, 299, 247, 223, 230, 246, 224, 229, 349, 223, 255, 223]
2
3Tokenization time: 0.0548 seconds
4
5Decoded output: ଆପଣ ଜଣେ ସହାୟକ ଏବଂ ଆପଣଙ୍କର ଭଲ ଐତିହାସିକ ଜ୍ଞାନ ଅଛି। ତୁମେ କେଉଁଠାରୁ ଆସିଛ? ମୁଁ ପୃଥିବୀରୁ ଆସିଛି
6
7Number of tokens: 56
8
9Comparison of Original and Decoded Text:
10
11Original Text: ଆପଣ ଜଣେ ସହାୟକ ଏବଂ ଆପଣଙ୍କର ଭଲ ଐତିହାସିକ ଜ୍ଞାନ ଅଛି। ତୁମେ କେଉଁଠାରୁ ଆସିଛ? ମୁଁ ପୃଥିବୀରୁ ଆସିଛି
12
13Decoded Text: ଆପଣ ଜଣେ ସହାୟକ ଏବଂ ଆପଣଙ୍କର ଭଲ ଐତିହାସିକ ଜ୍ଞାନ ଅଛି। ତୁମେ କେଉଁଠାରୁ ଆସିଛ? ମୁଁ ପୃଥିବୀରୁ ଆସିଛି| Tokenizer | Vocabulary Size | Tokenization Time (seconds) | Number of Tokens | Original Text | Decoded Output | Tokenized Output (Odia Characters) |
|---|---|---|---|---|---|---|
| shantipriya/minimind_odia_tokenizer | 150,000 | 0.563647 | 181 | ଓଡ଼ିଆ ଭାଷା ଏକ ଇଣ୍ଡୋ-ଆର୍ୟାନ୍ ଭାଷା... | ଓଡ଼ିଆ ଭାଷା ଏକ ଇଣ୍ଡୋ-ଆର୍ୟାନ୍ ଭାଷା... | ['Ċ', 'à¬ĵଡ', '଼ି', 'à¬Ĩ', 'Ġà¬Ń', ...] |
| ai4bharat/indic-bert | 200,000 | 2.369656 | 110 | ଓଡ଼ିଆ ଭାଷା ଏକ ଇଣ୍ଡୋ-ଆର୍ୟାନ୍ ଭାଷା... | ଓଡଆ ଭଷ ଏକ ଇଣଡ-ଆରୟନ ଭଷ... | ['[CLS]', '▁ଓ', 'ଡ', 'ଆ', '▁ଭ', 'ଷ', ...] |
| facebook/m2m100_418M | 128,104 | 0.358011 | 90 | ଓଡ଼ିଆ ଭାଷା ଏକ ଇଣ୍ଡୋ-ଆର୍ୟାନ୍ ଭାଷା... | ଓଡ଼ିଆ ଭାଷା ଏକ ଇଣ୍ଡୋ-ଆର୍ୟାନ୍ ଭାଷା... | ['en', '▁ଓଡ଼ିଆ', '▁ଭାଷ', 'ା', '▁ଏକ', ...] |
shantipriya/minimind_odia_tokenizer is the best choice due to its precision in tokenizing Odia words.facebook/m2m100_418M is ideal for handling many languages efficiently.ai4bharat/indic-bert offers robust support for multiple Indic languages but is slower in comparison to the other two.