A dataset of low-high quality Danish sentence pairs. For the bad sentence, we also include
a error type that describes the overall problem with the sentence.
The purpose of the dataset is to investigate linguistic quality in LLMs, similarly to a
linguistic acceptability dataset. However, the focus of this dataset goes broarder than
strictly acceptability, so that a "bad" sentence can be linguistically acceptable but
unnatural/disfluent.… See the full description on the dataset page:
https://huggingface.co/datasets/danish-foundation-models/linguistic-quality.