LongRCA Bench contains 1,140 observed, non-injected failed agent
trajectories from five task domains. Each trajectory has human annotations for
the responsible role, earliest decisive root-cause step, and a
trajectory-grounded rationale.
For the benchmark definition, annotation protocol, and evaluation results, see:
LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures
Source
Instances… See the full description on the dataset page:
https://huggingface.co/datasets/CLoud5-real/longrca-bench.