Combined 5 AI Conversational datasets, added a <|SYSTEM|> prompt for each, and broke the conversation down with <|USER|> and <|ASSISTANT|> tags.
You will need to add these tokens to your tokenizer to fully utilize this dataset: <|SYSTEM|> <|USER|> <|ASSISTANT|>
Collated dataset links:
Alpaca GPT-4
databricks-dolly-15k
Helpful and Harmless
Vicuna - English subset only
GPT4ALL-J
@misc{alpaca,
author = {Rohan Taori and Ishaan… See the full description on the dataset page:
https://huggingface.co/datasets/dmayhem93/ChatCombined.