This dataset contains 11 million raw Czech sentences.
11m-czech-sentences was created by filtering out sentences which contain at least one comma from the SYN2006PUB corpus.
This dataset is suitable for finetuning/training LLMs or other models on a small but very clean and correct representation of written language.
This dataset is merely a… See the full description on the dataset page:
https://huggingface.co/datasets/josefbednar/11m-czech-sentences.