LongProc (Long Procedural Generation) is a benchmark for evaluating long-context LLMs through long procedural generation tasks that require models to follow specified procedures and produce structured outputs. LongProc was accepted at COLM 2025.
LongProc consists of 6 tasks, each at up to 3 difficulty levels based on the expected output length (~0.5K, ~2K, ~8K… See the full description on the dataset page:
https://huggingface.co/datasets/PrincetonPLI/LongProc.