MiJaBench-Align is the companion response dataset for the Minority Jailbreaking Benchmark (MiJaBench).
While MiJaBench contains the adversarial prompts, MiJaBench-Align contains the model outputs generated when those prompts were evaluated across multiple large language models. Each record includes the original MiJaBench prompt metadata, the evaluated… See the full description on the dataset page:
https://huggingface.co/datasets/AKCIT/mijabench_align.