RULER long-context evaluation data, regenerated with the
Llama-2 tokenizer (NousResearch/Llama-2-7b-hf — the ungated mirror of
meta-llama/Llama-2-7b; LlamaTokenizer, SentencePiece, vocab_size=32000) so the labeled
context lengths are exact for Llama-2-family models instead of drifting, as they do when RULER
data tokenized for a different model (e.g. Qwen3) is fed to Llama-2.
Built with RULER's current ("binary-search") generators — so records carry… See the full description on the dataset page:
https://huggingface.co/datasets/tturing/ruler-500-llama2.