This dataset is a subset of aldea-ai/ruler-eval-data, containing only the 128K context length (131072 tokens). The full multi-length dataset also includes 1M and other lengths. Files are published under 131072/ (numeric token count) for compatibility with benchmark_ruler.py --context_length 131072, even though the source snapshot uses a 128k/ folder name.
Same as the upstream RULER on-disk layout, compatible with… See the full description on the dataset page:
https://huggingface.co/datasets/aldea-ai/ruler_eval_data_128k.