UltraChat-200k-enPurified is a highly curated, "prose-first" refinement of the mlabonne/ultrachat_200k_sft dataset.
The enPurified collection is built on a specific philosophy: Linguistic Specialization. While math and coding datasets are abundant, high-quality English prose often gets diluted by technical syntax or symbolic logic. This dataset isolates fluent, natural language to improve a model's conversational elegance and reasoning… See the full description on the dataset page:
https://huggingface.co/datasets/enPurified/ultrachat_200k_sft-enPurified-openai-messages.