RR (Circuit Breakers) attack completions with three-judge scores
This dataset bundles attack completions generated against
GraySwanAI/Llama-3-8B-Instruct-RR
(the "circuit breakers" defense), each scored by three independent judges:
local:strongreject (Lin et al., StrongREJECT classifier — most permissive)
local:harmbench (HarmBench classifier — middle)
local:gpt_oss (gpt-oss-safeguard-20b — strictest)
Headline finding: judges DISAGREE dramatically on certain… See the full description on the dataset page: https://huggingface.co/datasets/samuelsimko/rr-circuit-breakers-attack-completions.