GitHub repository:
https://github.com/XinXU-USTC/R2PE
Paper: Can We Verify Step by Step for Incorrect Answer Detection?
This is R2PE (Relation of Rationales and Performance Evaluation) Benchmark.
The aim is to explore the connection between the quality of reasoning chains and end-task performance.
We use CoT-SC to collect responses from 8 reasoning tasks spanning from 5 domains with various answer formats using 6… See the full description on the dataset page:
https://huggingface.co/datasets/xx18/R2PE.