Rubric-annotated agent trajectories for training and evaluating a rubric-conditioned
Process Reward Model (PRM) for data science agents.
Each example is a (task_description, trajectory, rubric_instruction) triple with a
4-class label produced by Claude Sonnet 4.6 as judge:
agent_success
The agent handled this aspect correctly… See the full description on the dataset page:
https://huggingface.co/datasets/atharva-naik-1/metadsprm-data.