EverythingLM V2 is a diverse instruct dataset consisting of 1k of human-assistant conversations. These sets were generated using principles from both evol-instruct and Orca. The dataset encompasses a wide array of topics and interactions.
All data in V2 is generated by GPT4
Higher quality dataset generation pipeline:
More humalike seed prompts
Fixed some bugs in the script
More diverse creative writing
More diverse seed prompts… See the full description on the dataset page:
https://huggingface.co/datasets/totally-not-an-llm/EverythingLM-data-V2.