ReTraceQA is a dataset designed to evaluate the reasoning traces of Small Language Models (SLMs) on commonsense reasoning tasks. It includes model-generated traces across four benchmark datasets: CommonsenseQA, OpenBookQA, QASC, and StrategyQA.
During the construction of ReTraceQA, only correct instances from the original benchmarks were retained, and erroneous instances were manually removed to ensure data quality.
Each item in… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/ReTraceQA.