The widespread availability of scientific documents in multiple languages, coupled with the development of automatic translation and editing tools, has created a demand for efficient methods that can detect plagiarism across different languages.
A dataset for cross-lingual plagiarism evaluation. Collection consists of a subset of Wikipedia articles on 4 languages (ru, hy, es, en). Query consists of wikipedia documents in… See the full description on the dataset page:
https://huggingface.co/datasets/AntiplagiatCompany/CL4Lang.