This dataset is created by sampling from multiple sources — ccdv/govreport-summarization test split,
MedRAG/pubmed train split,
osyvokon/pavlick-formality-scores test split,
and HuggingFaceGECLM/REDDIT_comments tifu split — to represent a diverse range of text snippets with varying levels of formality.
Each sample includes the fields "text", "type", "formality_label", and "formality_explanation".
For each sample, a large language model (LLM) DeepSeek-V3 was employed to assign a… See the full description on the dataset page:
https://huggingface.co/datasets/oishooo/formality_classification.