10,000 real-world dilemma / debate prompts used as the GRPO training prompts for a
dialectical-debate model. Each prompt is an open-ended question (advice dilemmas, opinion
debates, and general user requests) that the model is trained to answer by generating
multiple positions and against-claims in a structured "dialectical" format.
The prompts are drawn from public real-world sources: Reddit AITA
(r/AmItheAsshole), SHP (Stanford Human Preferences, a… See the full description on the dataset page:
https://huggingface.co/datasets/andreiski/dialectic-rl-questions-10k.