This dataset contains synthetic training examples for agentic RL.
Rather than simple prompt-response pairs, each sample is a self-contained agent task with a user goal, hidden scenario context, tool interfaces, operational constraints, adversarial pressure, and verifiable success criteria.
The data is generated or expanded by LLMs to create diverse workflows, tool ecosystems, and failure modes. As a result, the dataset is designed not just to train models to respond… See the full description on the dataset page:
https://huggingface.co/datasets/alibaba-pai/AgenticQwen-Data.