This dataset is a multilingual version of the original LAMBADA dataset dataset, which consists of a collection of narrative texts sharing the characteristic that human subjects are able to guess their last word if they are exposed to the whole text, but not if they only see the last sentence preceding the target word. To succeed on LAMBADA, computational models cannot simply rely on local context, but must be able to keep track of information in the broader… See the full description on the dataset page:
https://huggingface.co/datasets/Polygl0t/LAMBADA-poly.