This is the RL dataset for the paper: "ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL".
ReasonGen-R1 is a two-stage framework that imbues an autoregressive image generator with explicit text-based "thinking" skills via supervised fine-tuning (SFT) on a newly generated reasoning dataset of written rationales. It then refines its outputs using Group Relative Policy Optimization (GRPO). This dataset contains the model-crafted rationales paired with visual prompts… See the full description on the dataset page:
https://huggingface.co/datasets/Franklin0/ReasonGen-R1-RL-T2I-11k.