This dataset is a refined, high-precision monolingual collection derived from the DatarrX/myanmar-Wikipedia repository. It has been strictly filtered to ensure that every sentence contains exclusively Burmese characters, making it an ideal resource for language modeling, sequence-to-sequence tasks, and linguistic research where cross-lingual noise must be eliminated.
DatarrX (Burmese: ဒေတာအက်စ်) is a non-profit open-source… See the full description on the dataset page:
https://huggingface.co/datasets/DatarrX/burmese-only-wiki.