This contains the data created in the paper Fictitious Synthetic Data Can Improve LLM Factuality via Prerequisite Learning.
There are 4*2=8 configs in the dataset. The config names are formatted as {dataset}_{type}, where dataset is one of 'bio', 'medical', 'hotpotqa', and 'popqa', and type is 'knowledge' or 'skill'.
The data records with fake=True are the fictitious synthetic data created in this paper. The remaining data records are collected from existing datasets (except… See the full description on the dataset page:
https://huggingface.co/datasets/Shiyu-Lab/Prereq_Tune.