Dolma Reddit to Flashcards is a dataset of synthetically-generated QA items created on the basis of filtered Reddit data.
The creation of this dataset was motivated by the observation in Dolma (Soldaini et al. 2024) that the original Dolma Reddit data showed no benefit from inclusion of thread-level context over isolated submissions and comments, and that clean performance distinctions between tested Reddit versions were limited mainly to the HellaSWAG benchmark.
The… See the full description on the dataset page:
https://huggingface.co/datasets/allenai/dolma-reddit-to-flashcards-0625.