Status: experimental. Experiment-specific slice. Primary public dataset: scbe-aethermoore-training-data.
24,254 labeled prompts from 4 public prompt-injection datasets, each mapped through the Six Sacred Tongues bijective tokenizer from the SCBE-AETHERMOORE framework into a lossless per-prompt bit signature.
Stratified 70/15/15 train/val/test split by (source, label) so every source is represented in every split with its original label… See the full description on the dataset page:
https://huggingface.co/datasets/issdandavis/prompt-injection-bit-signatures.