β MarcoPolo Team β
Alibaba International Digital Commerce
Github π€ Hugging Face π Paper ποΈ Data
Given two aspects of search depth and width, existing benchmarks fall into four categories:
Low width, high depth benchmarks (e.g., GAIA, BrowseComp): focus on intricate deep reasoning over multi-hop retrieval for searching target answers
Low width, low depth⦠See the full description on the dataset page:
https://huggingface.co/datasets/ATH-MaaS/DeepWideSearch.