The pace of evolution of Large Language Models (LLMs) necessitates new approaches for rigorous and comprehensive evaluation. Traditional human annotation is increasingly impracticable due to the complexities and costs involved in generating high-quality, challenging problems. In this work, we introduce
CHASE, a unified framework to synthetically generate challenging problems using LLMs without human involvement. For a given task… See the full description on the dataset page:
https://huggingface.co/datasets/McGill-NLP/CHASE-Code.