25 human-verified multi-agent failure traces, each labelled with one of the 14 MAST
failure modes (Cemri et al., "Why Do Multi-Agent LLM Systems Fail?",
arXiv:2503.13657). Built as the ground-truth set for
evaluating the MAST classifier in
adk-agent-playground — i.e. for
Cohen's κ against an LLM judge, not for training.
id
stable id, e.g. gold-001
pipeline_name / agent_name
which agent produced… See the full description on the dataset page:
https://huggingface.co/datasets/barissozudogru/mast-failure-mode-gold.