Oolong-synth is a dataset from the paper Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities. See the paper for more details on the dataset construction.
To run the standard evaluation setting you will need:
input: context_window_text + "\n" + question (these are separated because the context window text can be cached for reuse across multiple input queries)
output: answer
UPDATE 6/20/2026: Corrected 14 instances, mostly for very-long-context temporal queries. Thanks to… See the full description on the dataset page:
https://huggingface.co/datasets/oolongbench/oolong-synth.