The Myanmar Tipitaka Dataset is a high-quality, structured collection of the Buddhist Pali Canon, transcribed in the Myanmar (Burmese) script. This dataset contains 462,504 paragraphs, covering the entire "Triple Basket" (Tipitaka) of Theravada Buddhism, including the original Mula (Canonical texts), Atthakatha (Commentaries), and Tika (Sub-commentaries).
This project was initiated by DatarrX to provide a clean… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/tipitaka-dataset.