This is a synthetic dataset generated using 📏 RULER: What’s the Real Context Size of Your Long-Context Language Models?.
It can be used to evaluate long-context language models with configurable sequence length and task complexity.
Currently, It includes 4 tasks from RULER:
QA2 (hotpotqa after adding distracting information)
Multi-hop Tracing: Variable Tracking (VT)
Aggregation: Common Words (CWE)
Multi-keys Needle-in-a-haystack (NIAH)
For each of the task, two target sequence lengths are… See the full description on the dataset page:
https://huggingface.co/datasets/rbiswasfc/ruler.