High-entropy prefix–suffix probes for measuring verbatim training-data leakage in
open-weight language models. This is the evaluation set used in Context as a Key:
Quantifying Verbatim Data Leakage Across Model Scale, Alignment, and Reasoning
Architectures (TSD 2026).
5,000 unique probes × 3 prefix lengths = 15,000 rows.
Each row gives you a prefix to feed a model and a target_suffix the model never
saw. If the… See the full description on the dataset page:
https://huggingface.co/datasets/fremy7/tsd2026-verbatim-leakage-probes.