This dataset is a diverse collection of parallel sentences in English and various other languages, sourced from multiple high-quality datasets.
Each sentence pair includes a semantic similarity score calculated using the Language-agnostic BERT Sentence Embedding (LaBSE) model,
along with additional quality metrics.
Machine Translation… See the full description on the dataset page:
https://huggingface.co/datasets/agentlans/en-translations.