Preprint is on arXiv and code for LLM evaluation is available on GitHub.
BABILong Leaderboard with top-performing long-context models.
BABILong is a novel generative benchmark for evaluating the performance of NLP models in
processing arbitrarily long documents with distributed facts.
It contains 11 configs, corresponding to different sequence lengths in tokens:… See the full description on the dataset page:
https://huggingface.co/datasets/RMT-team/babilong.