This dataset contains 21 many-shot ICL tasks that are designed to evaluate the long-context capability of LLMs, as introduced in the paper On Many-Shot In-Context Learning for Long-Context Evaluation. We categorize the tasks into similar-sample learning (SSL) and all-sample learning (ASL) groups.
SSL Tasks: banking77, dialogRE, TREC50, CLINC150, and BBH_geometric_shapes
ASL Tasks: GSM8K, MATH-algebra, MATH-counting_and_probability, MATH-geometry, MATH-number_theory, XLSUM, GPQA_cot… See the full description on the dataset page:
https://huggingface.co/datasets/launch/ManyICLBench.