HoloBench is a benchmark designed to evaluate the ability of long-context language models (LCLMs) to perform holistic reasoning over extended text contexts.
Unlike standard models that retrieve isolated information, HoloBench tests how well LCLMs handle complex reasoning tasks that require aggregating and synthesizing information across multiple documents or large text segments.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/megagonlabs/holobench.