A pre-processed version of the HybridDialogue dataset. The dataset was created as part of our work on Cerebras DocChat - a document-based conversational Q&A model. This dataset is intended to be used for training purposes, and so overlapping samples with the HybridDialogue test set in ChatRAG have been removed.
Each sample in this dataset contains a messages multi-turn conversation, a document which is a concatenated representation of relevant document(s), and… See the full description on the dataset page:
https://huggingface.co/datasets/cerebras/HybridDialogue.