LLM-judged root-cause labels for a 20K sample of agent tool-calling traces from
Agent-Ark/Toucan-1.5M,
using the B1–B8 agentic error taxonomy (as opposed to the hallucination-content
taxonomy used by the distill-reasoning/bert-spans datasets in this collection).
This is the first, flat-schema run — superseded by agentic-error-judge-v2
in this collection, which uses an updated multi-span-per-trace schema and covers
more traces. Kept here… See the full description on the dataset page:
https://huggingface.co/datasets/ssurface/agentic-error-judge-v1.