A synthetic instruction-tuning dataset for Python coding tasks, generated from the
verl open-source RL framework codebase.
Used to fine-tune Qwen/Qwen2.5-Coder-1.5B with LoRA, achieving +24 point absolute
improvement in pass@1 (0.565 → 0.804) after 3 epochs of training on a T4 GPU.
Category
Count
Description… See the full description on the dataset page:
https://huggingface.co/datasets/archit11/track_b_sft.