Multi-round questions and answers for randomly selected Wikipedia articles of varying lengths, in fastchat JSON format, generated by gpt-4-1106-preview. OpenAI terms apply.
This was designed to train a 32K context-length model. Check the total conversation lengths before using data items for training to ensure that they fit inside your target context window, and discard queries that don't fit.
Both the questions and answers were generated by GPT4, based on the document. Only information from… See the full description on the dataset page:
https://huggingface.co/datasets/grimulkan/wikipedia-document-question-answer.