A comprehensive dataset of MATH problem solutions generated by different language models with various sampling parameters.
Paper (arXiv) | GitHub Repository
This dataset contains code generation results from the MATH Dataset evaluated across multiple models. It was created to support the PIKA (Probe-Informed K-Aware Routing) project, which investigates how LLMs encode their own likelihood of success in their internal… See the full description on the dataset page:
https://huggingface.co/datasets/CoffeeGitta/pika-math-generations.