GRPO-fine-tuned Qwen/Qwen2.5-Coder-32B-Instruct for fuzzing harness generation. Trained on 10 C/C++ libraries (cJSON, curl, libjpeg, libtiff, libvpx, zlib, …) with four reward tasks: coverage, alignment, throughput, and stateful API interaction.
See the code repository for training scripts and the static-analysis backend.