Preliminary Historic Multilingual and Monolingual ByT5 Models. Following languages are currently covered:
More details can be found in
our GitHub repository.
In this experiment we sample 4B bytes (~4GB of text) from each corpora (and upsample Swedish and Finnish) and train for another epoch (2 epochs in total).
We use the official JAX/FLAX example in Hugging Face Transformers to pretrain a ByT5 model on a single v3-8 TPU.
Details about the training can be found
here.
Research supported with Cloud TPUs from Google's
TPU Research Cloud (TRC).
Many Thanks for providing access to the TPUs ❤️