Dataset Card for STaD Scaffolded Benchmarks
Dataset Summary
The STaD Scaffolded Benchmarks are diagnostic evaluation datasets introduced in the paper "STaD: Scaffolded Task Design for Identifying Compositional Skill Gaps in LLMs" (ACL Findings 2026). These benchmarks go beyond measuring whether a model answers correctly — they reveal where and why a model fails by breaking problems into sub-tasks and providing targeted scaffolding at specific reasoning steps.
Each… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/STaD.