C4 200M is a collection of 183,894,319 synthetic sentence pairs generated from the cleaned English portion of the C4 corpus for grammatical error correction (GEC).
This repository is a Parquet conversion of the original liweili/c4_200m dataset. The original dataset relied on a loading script, which is incompatible with recent versions of the 🤗 Datasets library. This version stores the data in Apache Parquet format, enabling efficient… See the full description on the dataset page:
https://huggingface.co/datasets/martinsr/c4_200m.