Data accompanying PEER: Unified Process–Outcome Reinforcement Learning for Structured Empathetic
Reasoning. Four subsets cover the full pipeline: reward-model supervision (SER), supervised
fine-tuning (SFT), reinforcement learning prompts (GRPO), and a safety probe suite.
Training and evaluation code: github.com/Yunxiao-Wang/PEER.
ser/
21,868 labels over 5,467… See the full description on the dataset page:
https://huggingface.co/datasets/YunxiaoWang/PEER.