Dataset Card for NLPre-PL – fairly divided version of NKJP1M
Dataset Summary
This is the official NLPre-PL dataset - a uniformly paragraph-level divided version of NKJP1M corpus – the 1-million token balanced subcorpus of the National Corpus of Polish (Narodowy Korpus Języka Polskiego)
The NLPre dataset aims at fairly dividing the paragraphs length-wise and topic-wise into train, development, and test sets. Thus, we ensure a similar number of segments
distribution per… See the full description on the dataset page: https://huggingface.co/datasets/ipipan/nlprepl.