This dataset accompanies the paper "Mind the Sim2Real Gap in User Simulation for Agentic Tasks" (arXiv:2603.11245).
TAU-USI is a human evaluation dataset for studying the sim-to-real gap in LLM-based user simulation for agentic tasks. As NLP evaluation shifts from static benchmarks to multi-turn interactive settings, LLM-based simulators have become widely used as user proxies. This dataset provides the… See the full description on the dataset page:
https://huggingface.co/datasets/cmu-lti/tau-usi.