VeriWeb: Verifiable Long-Chain Web Benchmark for Agentic Information-Seeking
[!NOTE]
This project was originally named VeriGUI. As our initial data collection focused on web-based tasks that primarily involve information-seeking rather than GUI interaction, we now define this part as the standalone VeriWeb benchmark, while desktop and other GUI-oriented scenarios will be released as a separate benchmark (in progress). We apologize for any resulting confusion.
Overview… See the full description on the dataset page: https://huggingface.co/datasets/2077AIDataFoundation/VeriWeb.