Agentic Reinforcement Learning (Agentic RL) has achieved notable success in enabling agents to perform complex reasoning and tool use.
However, most methods still relies on sparse outcome-based reward for training.
Such feedback fails to differentiate intermediate reasoning quality, leading to suboptimal training results.
In this paper, we introduce… See the full description on the dataset page:
https://huggingface.co/datasets/bunny127/Reagent-SFT-55.6K.