This dataset is a cleaned, English-only version of the Kyoto-Corpus dataset, filtered to remove all non-English conversations for supervised fine-tuning (SFT) of large language models.
The original Kyoto-Corpus contained multilingual conversations including Turkish, French, and other languages mixed with English content. This version provides a pure English subset suitable for training English-focused conversational AI models.… See the full description on the dataset page:
https://huggingface.co/datasets/Yxanul/yxanul-instruct-small.