width_20k.jsonl is constructed by us and is tailored for WideSearch-style tasks.
depth_20k.jsonl is sourced from ASearcher's training data.
hybrid_20k.jsonl is a balanced mixture of the two and serves as the core training set for our main training experience.
All three datasets are curated to 20,000 examples each.
Width Dataset… See the full description on the dataset page: https://huggingface.co/datasets/WideSeek-R1/WideSeek-R1-train-data.