This dataset was used to train CopyPasteLLM-L3-8B, presented in the paper Copy-Paste to Mitigate Large Language Model Hallucinations.
CopyPasteSeed365 is a high-quality seed dataset derived from three major RAG (Retrieval-Augmented Generation) benchmarks: PubMedQA, FaithEval, and RAGTruth. This dataset contains intermediate data from the DPO (Direct Preference Optimization) preparation pipeline, featuring complete responses and… See the full description on the dataset page:
https://huggingface.co/datasets/wingchiuloong/CopyPasteSeed365.