A ready-to-train, LettuceDetect-format
hallucination span dataset for token-classification (BIO-tagging) BERT models,
built from agent tool-calling traces in
Agent-Ark/Toucan-1.5M.
Balanced 50/50 between hallucinated and genuinely clean traces so a classifier
also learns what "no hallucination" looks like, not just where to find one.
toucan_train.jsonl — 4,654 rows total (split field: 4,190 train / 232 dev / 232 test).
Each row:
{
"uuid": "..."… See the full description on the dataset page:
https://huggingface.co/datasets/ssurface/toucan-train.