This dataset provides instruction-response pairs in the Tibetan language specifically designed for instruction fine-tuning of large language models.
The dataset contains 60,000 instruction-response pairs mix with Chinese and Tibetan, derived from:
Translated alpaca-gpt4 dataset (English instruction-following responses generated by GPT-4)
Chinese-Tibetan translation data
All content has undergone quality filtering to… See the full description on the dataset page:
https://huggingface.co/datasets/lightman7/tibetan-mix-instruction-tuning-60K.