This dataset is designed to assess text quality robustly across various domains for NLP and AI applications. It provides a composite quality score based on multiple classifiers, offering a more comprehensive evaluation of text quality beyond educational domains.
Dataset Details
Size: 100,000 sentences
Source: 20,000 sentences from each of 5 different datasets
allenai/c4
HuggingFaceFW/fineweb-edu… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/text-quality.