A high-quality instruction-tuning dataset derived from the official paper “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning”. This dataset contains curated reasoning-focused instruction–response pairs extracted from the training and evaluation protocols described in the paper, designed to support research on chain-of-thought (CoT) reasoning, reinforcement learning (RL), and model distillation.
Overview… See the full description on the dataset page: https://huggingface.co/datasets/amishor/reinforce-learning-grpo.