PandaBench is a comprehensive benchmark for evaluating Large Language Model (LLM) safety, focusing on jailbreak attacks, defense mechanisms, and evaluation methodologies.
The PandaGuard framework architecture illustrating the end-to-end pipeline for LLM safety evaluation. The system connects three key components: Attackers, Defenders, and Judges.
This repository contains the benchmark results from extensive evaluations of various… See the full description on the dataset page:
https://huggingface.co/datasets/Beijing-AISI/panda-bench.