A high-quality, large-scale Burmese (Myanmar) language dataset extracted and refined from the official Burmese Wikisource. This dataset is specifically designed for Natural Language Processing (NLP) tasks, including language modeling, spell checking, and formal text analysis.
Organization: DatarrX
Maintainer: Khant Sint Heinn(Kalix Louis)
Source: Burmese Wikisource (my.wikisource.org)
Total Rows: 1,352,972
Total Rows: 1,352,972… See the full description on the dataset page:
https://huggingface.co/datasets/DatarrX/mywikisource.