This dataset contains frozen inputs for the RULER benchmark.
Benchmark: RULER
Sequence length: 4,096 tokens
Tokenizer: Qwen/Qwen3.5-9B
Tokenizer revision: c202236
lm-eval version: 0.4.12
Task configurations: 13
Samples per configuration: 500
Deterministic generation: Yes. Each configuration resets Python, NumPy, and task random state to seed 42.
niah_single_1
niah_single_2
niah_single_3… See the full description on the dataset page:
https://huggingface.co/datasets/khashazad/amc-ruler-qwen35-4k.