This dataset is a high-quality, large-scale Burmese language corpus curated from a diverse range of sources, including classical literature, news, encyclopedic content, and conversational data. It has been rigorously cleaned to ensure linguistic integrity and is suitable for various NLP tasks such as Language Modeling, Tokenizer Training, and Text Classification.
The kalixlouiis/raw-data corpus is a synthesized… See the full description on the dataset page:
https://huggingface.co/datasets/kalixlouiis/raw-data.