An open, fully-local benchmark of failure modes in multi-agent LLM systems —
750 execution traces from three frameworks (CrewAI, AutoGen,
LangGraph) run on a local Llama 3.1 8B model (via Ollama) at $0 API cost.
The canonical, human-validated taxonomy of multi-agent LLM failures is MAST —
Cemri et al., "Why Do Multi-Agent LLM Systems Fail?" (NeurIPS 2025). AgentFailDB does
not supersede or replace it; it is a fully-local… See the full description on the dataset page:
https://huggingface.co/datasets/Jeevaa79/agentfaildb.