This is just the original ultrachat 200k dataset converted to sharegpt format.
This is a heavily filtered version of the UltraChat dataset and was used to train Zephyr-7B-β, a state of the art 7b chat model.
The original datasets consists of 1.4M dialogues generated by ChatGPT and spanning a wide range of topics. To create UltraChat 200k, we applied the following logic:
Selection of a subset of data for faster… See the full description on the dataset page:
https://huggingface.co/datasets/abhinand/ultrachat_200k_sharegpt.