This repository contains the dataset for the paper CyberThreat-Eval: Can Large Language Models Automate Real-World Threat Research? (published in TMLR).
CyberThreat-Eval is an expert-annotated benchmark collected from the daily Cyber Threat Intelligence (CTI) workflow of a world-leading company. It assesses Large Language Models (LLMs) on practical tasks across three essential stages of threat research.
Stage 1: Triage —… See the full description on the dataset page:
https://huggingface.co/datasets/xse/CyberThreat-Eval.