ShellRisk-Bench is a reproducible benchmark for context-free binary risk
classification of individual shell-command submissions. It asks whether a
command poses meaningful cyber or system risk when evaluated without task,
user, or session context.
Release: The v0.1 Parquet train and test splits are publicly available
through Dataset Viewer and load_dataset().
The benchmark contains a deterministic train split of 16,772 rows and test
split of 4,194 rows. The test… See the full description on the dataset page:
https://huggingface.co/datasets/kontext-security/ShellRisk-Bench.