The dataset contains parallel corpora, featuring pairs in Russian and English.
Initially, the original OPUS PHP v1 (en&ru) dataset was intended to be used for training the enbeddrus project. However, due to its poor quality and high level of noise, it was decided to use only the English corpus from this dataset. The English text was then automatically translated using LibreTranslate, followed by manual translation and quality improvement using… See the full description on the dataset page:
https://huggingface.co/datasets/evilfreelancer/opus-php-en-ru-cleaned.