ARC is a dataset that tests the level of understanding of science at approximately grade-school level.
We focus specifically on the 'Challenge' subsection of ARC, the more difficult of the two subsections, which has been widely adopted as a measure of LLM reasoning and world understanding.
We create a paired preference-ranked dataset from the train split of ARC-Challenge.
The dataset is partitioned into questions which we take as our prompts x… See the full description on the dataset page:
https://huggingface.co/datasets/abacusai/ARC_DPO_FewShot.