Measuring whether an agent can hold a workstream together across many turns and many systems: long-horizon, multi-tool reconciliation under adversarial conditions.
Layout · Performance · Dataset · Environment · Grading · Adversarial design · Task structure · Trajectories · Verification
Dōjigiri measures real-world assistant competence over a long horizon, not isolated tool calling, retrieval, or single-turn… See the full description on the dataset page:
https://huggingface.co/datasets/ethara/dojigiri-samples.