Dataset Card for ProcBench
Dataset Overview
Dataset Description
ProcBench is a benchmark designed to evaluate the multi-step reasoning abilities of large language models (LLMs). It focuses on instruction followability, requiring models to solve problems by following explicit, step-by-step procedures. The tasks included in this dataset do not require complex implicit knowledge but emphasize strict adherence to provided instructions. The dataset evaluates model… See the full description on the dataset page: https://huggingface.co/datasets/ifujisawa/procbench.