Token usage count per language when tokenizing the "bigcode/stack-dedup-alt-comments" dataset with the santacoder tokenizer. There are less tokens than in the tokenizer because of vocabulary mismatch between the datasets used to train the tokenizer and the ones that ended up being used to train the model.
More Information needed