This is a set of wordlists extracted from the GlotLID-corpus for high-precision filtering of FineWeb2.
The recommended way to download the data is via git clone:
git clone
https://huggingface.co/datasets/cis-lmu/glotlid-wordlists
For details on the filtering method, please refer to the FineWeb2 paper.
Each word listed occurs significantly more often in its own language dataset than in any… See the full description on the dataset page:
https://huggingface.co/datasets/cis-lmu/glotlid-wordlists.