Dataset automatically created during the evaluation run of model notbdq/Qwen2.5-14B-Instruct-1M-GRPO-Reasoning
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page:
https://huggingface.co/datasets/open-llm-leaderboard/notbdq__Qwen2.5-14B-Instruct-1M-GRPO-Reasoning-details.