Training data for offline GRPO (Group Relative Policy Optimization) on IPDA debate generation.
Dataset Structure
Each row represents one pipeline call with 4 response variants:
RESPONSE_1_* through RESPONSE_4_*: Different generations at varying temperatures
*_SCORE: Quality score (0.0-1.0) from Haiku evaluator
chosen_index: Index of highest-scoring response
rejected_index: Index of lowest-scoring response
Statistics… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/debate-multi-trial-grpo.