A gated research dataset of AI agent trajectories, rendered as markdown, used in a
human-subjects study on the oversight of AI agents. The set contains 160 stimuli
spanning two task domains: customer service and software engineering.
Access is gated while the associated work is in preparation. The public card is kept
minimal on purpose. Everything you need to use the data (the schema, the data
dictionary, the intended analysis, and the… See the full description on the dataset page:
https://huggingface.co/datasets/evijit/HumanOversightBench.