This dataset contains 242,198 revision histories from Wikipedia, WikiNews, and Arxiv, capturing sentence-level changes across document versions. This enables the study of how documents evolve over time, particularly at the granular level of sentences.
The Wikipedia and WikiNews revision data was collected using this custom web crawler, while the Arxiv data was gathered using the crawler from the IteraTeR project.