Two-stage training data for the WebArbiter process reward model
Published at ICLR 2026
Paper | Code | Website | Collection | Demo
This repository contains the training data for WebArbiter, a principle-guided reasoning Process Reward Model (PRM) for web agents. We build on the WebPRM Collection (Chae et al., 2025), which comprises ~30k step-level preference pairs drawn from the Mind2Web environment. WebArbiter is trained via a… See the full description on the dataset page:
https://huggingface.co/datasets/ZYao720/WebArbiter-Data.