NovelPrompts is an English safety dataset (194 prompts + completions)
built to test LLM judges' behaviour when evaluating the safety of prompts.
The prompts are designed such that safety can only be assessed if you understand
a novel concept: a recent event, a new word, or a new meaning of an existing word
(e.g. slang). A concept is considered novel if it appeared after July 2024.
This dataset is intended for evaluating LLM-as-judge safety evaluators, especially… See the full description on the dataset page:
https://huggingface.co/datasets/anissa218/novelprompts.