Wikipedia - 20000 rows
PersonaChat truecased - 20000 rows
Synthetic edge case data - 5000 rows
Synthetic quoted text data - 5000 rows
The synthetic data has been generated using GPT-5.3 models.
The other data was sourced from the original Hugging Face sources.
This dataset can be used to train text normalizers that convert badly formatted English into correct English.… See the full description on the dataset page:
https://huggingface.co/datasets/lunahr/normalization-data-mixed.