Sentence-level evaluation of the comma placement abilities of large language models in Norwegian sentences.
The NCB corpus (version 0.1) is a collection of 840 human-written Norwegian sentence pairs. The sentences are manually collected from publicly available sources such as articles and governmental reports. The sentences aim to be representative of Norwegian non-fiction, in particular governmental… See the full description on the dataset page:
https://huggingface.co/datasets/hcfa/ncb.