This dataset is the Tibetan-English sentence pairs from the NLLB dataset (see block quote below).
This dataset was found on, and downloaded from, OPUS.
The dataset was re-formatted for Hugging Face using the code found here.
The topic labels were generated with easy_text_clustering using the code found here.
This dataset was created based on metadata for mined bitext released by Meta AI. It contains bitext for 148 English-centric and 1465 non-English-centric language pairs using the stopes… See the full description on the dataset page:
https://huggingface.co/datasets/billingsmoore/NLLB-bo-en.