Document extraction is a multi-step process, and an end-to-end accuracy number hides where it broke. ProcessBench evaluates reasoning step by step, which is the shape of evaluation we want for our own pipelines.
We use it for: benchmarking step-level error localization - studying how failures propagate through a chain rather than just counting final-answer mistakes.