Wikipedia dataset containing cleaned articles of Polish language.
The dataset has been built from the Wikipedia dump (
https://dumps.wikimedia.org/)
using the OLM Project.
Each example contains the content of one full Wikipedia article with cleaning to strip
markdown and unwanted sections (references, etc.).
Most of Wikipedia's text and many of its images are co-licensed under the
Creative Commons… See the full description on the dataset page:
https://huggingface.co/datasets/chrisociepa/wikipedia-pl-20230401.