RL (GRPO) dataset for the LoRAcle pipeline. Paired with ceselder/loracle-ia-warmstart-v5.
400 LoRAs, 1 row per LoRA:
200 from the warmstart_v5 pool (already-seen) — RL fine-tunes on familiar LoRAs
200 from the held-from-warmstart pool — RL must generalize to unseen LoRAs
Same as loracle-ia-warmstart-v5: lora_id, source, qa_type, question, answer, ground_truth, category.
203 rows from the original… See the full description on the dataset page:
https://huggingface.co/datasets/ceselder/loracle-ia-RL-v5.