SEA-PILE v2 is a large, multilingual language modelling dataset of 120 billion tokens, sourced from a diverse array of web content.
Languages supported: Vietnamese, Bahasa Indonesia, Tamil, Malay, Thai, Tagalog, Khmer, Lao, Burmese
The total number of tokens in the dataset has been calculated using the Gemma3 tokenizer
Language
ISO 639-1 Code
Total Number of Tokens (Billions)
Percentage
Vietnamese
vi
51.4
42.13%… See the full description on the dataset page:
https://huggingface.co/datasets/aisingapore/SEA-PILE-v2.