OpenHermes-2.5-Code-290k
This dataset is amalgamation of two datasets. I have used OpenHermes-2.5 a super quality dataset made avaliable by teknium.
Other datset is my own Code-290k-ShareGPT.
This dataset is in Vicuna/ShareGPT format. There are around 1.29 million set of conversations.
I have cleaned the dataset provided by Teknium and removed metadata such as "source" & "category" etc.
This dataset has primarily synthetically generated instruction and chat samples.
This dataset is very… See the full description on the dataset page:
https://huggingface.co/datasets/ajibawa-2023/OpenHermes-2.5-Code-290k.