The WebInstruct-Verified-Processed dataset is used in the paper Characterizing, Evaluating, and Optimizing Complex Reasoning. It is a processed version of WebInstruct-verified, formatted for RL training with verifiable rewards. In the paper, this dataset is used as the RL prompt/data source for TRM-guided reinforcement learning optimization, where models are trained with rule-based verifiers and an auxiliary Thinking Reward Model (TRM) signal.… See the full description on the dataset page:
https://huggingface.co/datasets/zzzhr97/WebInstruct-Verified-Processed.