This dataset was presented in the paper CHALIS: A Challenge Dataset for Language Identification in Difficult Scenarios.
CHALIS is a multilabel dataset composed of several sections aimed at testing language identification abilities in difficult scenarios.
Main contribution comes in the form of gathering and classifying by a human expert of sentences belonging to four language pairs (Spanish - Catalan, Portuguese - Galician… See the full description on the dataset page:
https://huggingface.co/datasets/michal-tichy/CHALIS.