This dataset contains English-Polish parallel sentence pairs with precomputed English sentence embeddings from Qwen/Qwen3-Embedding-4B.
The source data was scored with google/metricx-24-hybrid-large-v2p6. Only examples with metricx_pred <= 7.0 were retained, so this is a quality-filtered subset rather than the full original parallel corpus.