This dataset is built from the open source data accompanying "An Open Dataset and Model for Language Identification" (Burchell et al., 2023)
The repository containing the actual data can be found here :
https://github.com/laurieburchell/open-lid-dataset.
The license for this recreation itself follows the original upstream dataset as GPLv3+.
However, individual datasets within it follow each of their own licenses.
The "src" column lists the sources. "lang" column lists the language code in… See the full description on the dataset page:
https://huggingface.co/datasets/hac541309/open-lid-dataset.