SHARP is a corpus for step-level hallucination detection in mathematical reasoning. Open 7B models
were asked to solve GSM8K and MATH problems, and a large teacher model then marked the exact fragments of
each solution that are wrong or unsupported. The result is a step-labelled corpus that can be used to train
Process Reward Models (PRMs), to benchmark error localization, or to study where small models go off the rails.
The corpus is published… See the full description on the dataset page:
https://huggingface.co/datasets/ZaandaTeika/SHARP.