The quality of Supervised Fine-Tuning (SFT) data plays a critical role in enhancing the conversational capabilities of Large Language Models (LLMs).
However, as LLMs become more advanced,
the availability of high-quality human-annotated SFT data has become a significant bottleneck,
necessitating a greater reliance on synthetic training data.
In this work, we introduce… See the full description on the dataset page:
https://huggingface.co/datasets/internlm/Condor-SFT-20K.