myX-Mega-Corpus is a comprehensive, large-scale Burmese language dataset curated by DatarrX. It consists of approximately 16 million rows of cleaned and shuffled Burmese text. While originally designed for training the myX-Semantic word embedding model, this corpus is highly versatile and can be used for any NLP task, including LLM fine-tuning, sentiment analysis, and machine translation.
📚 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/myX-Mega-Corpus.