This dataset is a subset of the
https://huggingface.co/datasets/stingning/ultrachat.
This dataset focuses on the question answering task on an existing context, using a simple keyword filter (any question containing one of these keywords: passage, article, context). I also extract only the first round of conversation and convert it to the familiar alpaca format, and further filter so that the dataset only contain long input (which means… See the full description on the dataset page:
https://huggingface.co/datasets/nguyenthanhdo/ultrachat-aem-alpaca-v1.0.