Training data for sampling-based watermark distillation using the KGW k=0,γ=0.25,δ=2k=0, \gamma=0.25, \delta=2k=0,γ=0.25,δ=2 watermarking strategy in the paper On the Learnability of Watermarks for Language Models. Llama 2 7B with decoding-based watermarking was used to generate 640,000 watermarked samples, each 256 tokens long. Each sample is prompted with 50-token prefixes from OpenWebText (prompts not included… See the full description on the dataset page:
https://huggingface.co/datasets/cygu/sampling-distill-train-data-kgw-k0-gamma0.25-delta2.