TIFD (Tibetan Instruction-Following Dataset) is a specialized instruction dataset for large language models supervised fine-tuning. The dataset contains 11,535 high-quality Tibetan instructions with four attributes: unique identifier, instruction, input, and output.
Scale: 11,535 high-quality Tibetan instruction data
Format: JSON format with four fields: id, instruction, input, output
Source: Generated by… See the full description on the dataset page:
https://huggingface.co/datasets/CMLI-NLP/TIFD.