A comprehensive, unified benchmark dataset for evaluating Large Language Models (LLMs) on cybersecurity tasks. This dataset combines 10 security benchmarks into a standardized format with 18,059 examples across 5 task types.
This dataset consolidates multiple security-focused benchmarks into a single, easy-to-use format for comprehensive LLM evaluation across various cybersecurity domains:
Total Examples: 18,059
Total⦠See the full description on the dataset page:
https://huggingface.co/datasets/tuandunghcmut/combine-llm-security-benchmark.