UserIntentBench is a benchmark for measuring whether LLM and VLM agents can recover, track, and re-align with user intent over long-horizon agentic tasks. Real user intent is rarely fully specified and rarely fixed: it is latent (only partially expressed at the start) and shifting (refined, redirected, or replaced mid-session). The benchmark represents the user's full intent as a structured, web-grounded intent graph that the evaluation harness knows in full but… See the full description on the dataset page:
https://huggingface.co/datasets/JingmingChen/UserIntentBench.