An ablation study of reinforcement-learning (RL) fine-tuning for agentic software-engineering (SWE) models. Starting from an 8B SFT model, we fine-tune with RL across ~20 configurations — varying the objective, loss normalization, sampling, and training dataset — and evaluate each on agentic SWE benchmarks.
RL reliably and substantially improves agentic SWE performance, and the improvement is… See the full description on the dataset page:
https://huggingface.co/datasets/penfever/ablation_exploration_in_rl.