rLLM is an open-source project to fully democratize reinforcement learning (RL) for LLMs and reproduce DeepSeek R1 and OpenAI O1/O3 at scale on real tasks. For all releases, we open source all our efforts here-including training scripts (including hyperparameters), models, systems, dataset, and logs.
DeepCoder's LiveCodeBench (LCB) score as training progresses. At step 180, context⦠See the full description on the dataset page:
https://huggingface.co/datasets/Ashenone3/rllm_temperature.