[!NOTE]
For more information on the ORBIT dataset, go check out the preprint available at arxiv.org/abs/2604.01195.
ORBITis a reasoning-intensive synthetic dataset with complex queries used for training search agents, generated without relying on any paid API services or manual annotation.
Training data for deep search — tasks requiring multi-step retrieval and reasoning over the web — is scarce.… See the full description on the dataset page:
https://huggingface.co/datasets/orbit-ai/orbit-20k.