OpenSkillRisk is a benchmark for evaluating whether CLI agents can safely handle risky third-party skills in open skill ecosystems.
This dataset contains the skill files and task specifications used by OpenSkillRisk. The code for building sandboxes and running evaluations is released separately at:
Repository: Miaow-Lab/OpenSkillRisk
Paper: OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party… See the full description on the dataset page: https://huggingface.co/datasets/Miaow-Lab/OpenSkillRisk.