This is a synthetic dataset generated with the CRAFT framework proposed in the paper CRAFT Your Dataset: Task-Specific Synthetic Data Generation Through Corpus Retrieval and Augmentation.
The correctness of the data has not been verified in detail, but training on this data and evaluating on human-curated commonsense question-answering data proved highly beneficial.
4 synthetic dataset sizes (S, M, L, XL) are available, and training on them yields consistent… See the full description on the dataset page:
https://huggingface.co/datasets/ingoziegler/CRAFT-CommonSenseQA.