184 annotated failure tasks collected from
Algorithm-generated agentic systems built using CaptainAgent,
Hand-crafted systems such as Magnetic-One.
Fine-grained annotations for each failure, including:
The failure-responsible agent (who failed),
The decisive error step (when the critical error occurred),
A natural language explanation of the failure.
The dataset covers a wide range of realistic multi-agent scenarios… See the full description on the dataset page:
https://huggingface.co/datasets/Kevin355/Who_and_When.