This is the LAMBADA test split modified for bidirectional language models (for example BERT). The original is appended by punctuation symbols (for example ."), as predicted by GPT-2 (small). The original is the LAMBADA test split as pre-processed by OpenAI,
LAMBADA is used to evaluate the capabilities of computational models for text understanding by means of a word prediction task. LAMBADA is a collection of narrative texts sharing the characteristic that human… See the full description on the dataset page:
https://huggingface.co/datasets/ltg/lambada-context.