This dataset is AFAIK (12 january 2024) the biggest ready to use open source dataset to finetune LLMs. It contains more than 3.8 million chat samples.
Its a collection of multiple different datasets. Some of them have been built using GPT4 or using scraped data. Here is the list:
teknium1_GPTeacher_codegen… See the full description on the dataset page:
https://huggingface.co/datasets/guigux/hulk_dataset_0.1.