Per-model classifications and summary statistics for the 13-model behavioral benchmark of Paper 2, "Is Code Safety a Separate Direction? A 13-Model Behavioral Benchmark of Malicious-Code Refusal in Coding-Specialized LLMs" (Young & Moody, 2026).
This dataset contains only judge labels and aggregate statistics. Prompt text is not redistributed here (it lives in the companion Paper 1 prompt-bank), and raw model responses are not redistributed… See the full description on the dataset page:
https://huggingface.co/datasets/richardyoung/code-safety-benchmark-results.