This dataset is the Tibetan-English sentence pairs from the Tatoeba dataset. It was found on, and downloaded from, OPUS. See more below.
This data was reformatted for Hugging Face from the moses format using the code found here.
The topic labels were generated with easy_text_clustering using the code found here.
Corpus Name: Tatoeba
Package: Tatoeba.bo-en in Moses format
Website:
http://opus.nlpl.eu/Tatoeba-v2023-04-12.php
Release: v2023-04-12
Release date: Wed Apr 12 23:14:42… See the full description on the dataset page:
https://huggingface.co/datasets/billingsmoore/Tatoeba-bo-en.