π WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation
Existing web agent evaluation suffers from three key limitations: (1) insufficient benchmark coverage β offline benchmarks lack real-world fidelity, while online benchmarks remain limited in website scale, domain diversity, and intent variety, leading to biased and overly optimistic assessments; (2) unscalable evaluationβ¦ See the full description on the dataset page:
https://huggingface.co/datasets/Mininglamp-2718/WebRetriever.