SSPO
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents
📄 arXiv •
💻 Code •
🤗 Dataset
🌟Overview
Deep search agents operate over trajectories spanning dozens of information-seeking steps, but standard reinforcement learning provides only a single outcome reward for the entire trajectory. This sparse signal makes it difficult to determine which intermediate reasoning and tool-use actions should be reinforced… See the full description on the dataset page: https://huggingface.co/datasets/WaitHZ/SSPO-data.