This dataset consists of a single csv with 10,000,000 rows which are pairs of sentences or phrases. The first member of each pair is a sentence or phrase in Classical Tibetan. The second member is the English translation of the first.
The pairs are pulled from texts sourced from Lotsawa House and are offered under the same license as the original texts from which they are sourced.
This data was scraped, cleaned, and formatted programmatically. Because of the difficulty in assembling data in… See the full description on the dataset page:
https://huggingface.co/datasets/billingsmoore/phonetic-tibetan-to-english-translation-pairs.