The MCWC is a curated multilingual corpus of constitutional texts from 191 countries, including both current and historical versions. The dataset provides aligned constitutional content in English, Arabic, and Spanish, enabling comparative legal analysis and multilingual NLP research.
This CSV version is a cleaned, structured, and sentence-aligned representation of the corpus, suitable for machine translation, information… See the full description on the dataset page:
https://huggingface.co/datasets/drelhaj/MCWC.