LooGLE is a comprehensive evaluation benchmark for LLM long context understanding which contains up-to-date (all after 2022) and extremely long realistic documents (over 24k tokens per document, many of which exceed 100k words) and 6,000 newly generated questions spanning diverse domains and categories. Details statistics of our dataset can be seen in the table below.
Short and long dependency tasks LooGLE is composed of 7 major tasks to evaluate LLMs' ability to… See the full description on the dataset page:
https://huggingface.co/datasets/bigai-nlco/LooGLE.