This dataset is a collection of malaysian texts in the Malay, English, Chinese, and Tamil languages, gathered by Malaysia AI volunteers through web crawling of malaysian websites.
The dataset amounts to approximately 250 GB of text data, and has undergone deduplication process.
To learn more about this project,
https://github.com/users/huseinzol05/projects/1/views/1
We no longer update the project.