Subset of 12.500 Wikipedia files extracted from the Wiki727k Text Segmentation dataset for my master thesis.
Distribution of Wikipedia Content is managed under the Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) License.
Text Segmentation is the task of splitting a text into semantic coherent paragraphs.
Each file has a number of paragraphs, divided by '========,,.