Status: draft. This is the audited validation-scale release (517 questions) of a
dataset that labels open-domain questions by whether a mid-size LLM needs web
search to answer them correctly. A train-split scale-up (~509k questions) is in
progress and will be released separately.
Built as part of a semester project on a web-search MCP service and shared
Elasticsearch indexing infrastructure for the Swiss AI initiative… See the full description on the dataset page:
https://huggingface.co/datasets/helkhour/web-search-router-labels.