JO-Bench is a curated benchmark of harmful prompts used to evaluate LLM safety, as introduced in the paper:
Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem (MLSys 2026).
This dataset combines two existing public benchmarks to create a specialized evaluation set of 128 prompts:
JailbreakBench (Chao et al., 2024): 100 samples.
HarmBench (Chemical & Biological category, Mazeika et al.… See the full description on the dataset page: https://huggingface.co/datasets/shuyilin/JO-Bench.