This repository contains the data used in our paper Broken-Token: Filtering Obfuscated Prompts by Counting Characters-Per-Token.
While we release this dataset under the CC-by-NC-SA-4.0 license, the dataset was constructed and built on other datasets, each with its own license, as mentioned below:
SoftAge-AI/prompt-eng_dataset, MIT License
Aiden07/dota2_instruct_prompt, MIT License
hassanjbara/ghostbuster-prompts, MIT License… See the full description on the dataset page:
https://huggingface.co/datasets/jfrog/obfuscation-identification.