Representative subsample of the test split for reviewer inspection.
One counterfactual and one non-counterfactual sample were selected per task type from randomly presented candidates drawn from the test split. Samples were accepted or skipped based on whether they passed the same validation criteria used during full dataset annotation, not on answer quality or difficulty.
qa.parquet - QA samples (17 rows: 9 non-CF + 8 CF across 9 task types)… See the full description on the dataset page:
https://huggingface.co/datasets/goldsmith-herbal-clay/360qvr-subsample.