Strata-Sword: A Hierarchical Safety Evaluation towards LLMs based on Reasoning Complexity of Jailbreak Instructions
Strata-Sword Strata-Sword is a multi-level safety evaluation benchmark proposed by Alibaba AAIG team. It aims to more comprehensively assess models' safety capabilities when facing jailbreak instructions of varying reasoning complexity, helping model developers better understand each model's safety boundaries.
🧩 Our Approach — Strata-Sword… See the full description on the dataset page: https://huggingface.co/datasets/OysterAI/Strata-Sword.