29,695,269 documents across 5,954 locales (language–script–region combinations; 297 languages, 322 language–script pairs), each annotated with topic assignments and a cultural-taxonomy profile. Text derives from FineWeb and FineWeb-2 via the FineWeb-CLaR region-attribution pipeline.
Per locale (up to 10,000 quality-filtered documents): BGE-M3 embeddings → FASTopic (500 topics… See the full description on the dataset page:
https://huggingface.co/datasets/Yusser/FineWeb-CLaR-culture.