This upload contains a combined bilingual datasets from various reliable sources.
Each source was downloaded and had english <--> polish samples extracted and saved to individual files.
Each .json file contains a list of objcects with keys "en" and "pl".
[
{"en":"text", "pl":"tekst"},
]
This is a convenient way to grab all the top-quality bilingual datasets for this pair.
Below is the list of actual datasets used for extraction… See the full description on the dataset page:
https://huggingface.co/datasets/laelhalawani/en_pl_collection.