MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
MoNaCo is a benchmark of 1,315 human-written time-consuming questions that require retrieval, filtering and aggregation across text and tables --- with an average of 43.3 distinct documents per question!
The broad scope of MoNaCo questions makes it ideal as an LLM benchmark for at least five different settings:
Factuality: Evaluating models’ parametric… See the full description on the dataset page:
https://huggingface.co/datasets/allenai/MoNaCo_Benchmark.