Info:
Translated on Russian by Google Translate
Source: xTRam1/safe-guard-prompt-injection
Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns
Size: 1,000 prompts (500 safe / 500 unsafe)
Columns:
text - original prompt
label - 0: safe, 1: unsafe
translation - prompt on Russian translated by Google Translate
score_ru_google - cosine similarity score with codebook
More information in paper:… See the full description on the dataset page:
https://huggingface.co/datasets/shalanova/benchmark-2-russian-gt.