A curated dataset for reinforcement learning (RL) training within the MiroRL framework.
Source: Provided by MiroMind AI as part of the MiroRL project.
Format & Size: Contains ~13.1k examples in Parquet format for efficient loading and processing.
License: Released under CC-BY-NC-4.0 for non-commercial use.
Purpose: Designed to serve as high-quality input for RL fine-tuning in the MiroRL pipeline.
Each record… See the full description on the dataset page:
https://huggingface.co/datasets/miromind-ai/MiroRL-GenQA.