This dataset is a deduplicated subset of ARC-Easy, as used in the paper How Much Can We Forget about Data Contamination?. The deduplication was performed using this script.
The data fields are the same as in
https://huggingface.co/datasets/allenai/ai2_arc, with the additional "split-id" column that can be used to partition the benchmark questions into different subsets.
The dataset can be used as a plug-in replacement for ARC-Easy if you want to work with the deduplicated benchmark questions.… See the full description on the dataset page:
https://huggingface.co/datasets/sbordt/forgetting-contamination-arc-easy.