This dataset contains 5.6M culturally tagged samples as presented in the paper The Culture Funnel: You Can't Align What isn't in the Data.
This dataset is designed to help researchers study and mitigate the "cultural data funnel" in Large Language Model (LLM) pipelines. We use a multidimensional tagging framework to identify cultural signals, domains, geographic locations, and task… See the full description on the dataset page:
https://huggingface.co/datasets/CohereLabs/CultureMarkers.