This benchmark checks whether a language model can recognise when a conversation is itself part of an evaluation rather than normal, real-world usage.
For full details see the accompanying paper: “Large Language Models Often Know When They Are Being Evaluated” (arXiv: 2505.23836) –
https://arxiv.org/abs/2505.23836.
true – the conversation was written for evaluation (e.g. prompt‐injection tests… See the full description on the dataset page:
https://huggingface.co/datasets/jjpn2/eval_awareness.