Repackaged release of the Complex Word Identification (CWI) Shared Task 2018 corpus (BEA-13 @ NAACL 2018), converted from the original TSV distribution into parquet with a unified schema across all four languages.
The shared task asks systems to predict whether a target word/phrase in context would be hard to understand for non-native speakers, children, or readers with language disabilities. It supports both a binary classification… See the full description on the dataset page:
https://huggingface.co/datasets/alvations/complex-word-id-2018.