BPE tokeniser with vocabulary size 40000, trained on the first 30 million examples in English
OSCAR.
BPE trainer implementation (note that these are long deprecated and you should use
TkTkT):
Preprocessor: see the appendix of the
BPE-knockout paper.