BiST is a large-scale bilingual translation dataset, with "BiST" standing for Bilingual Synthetic Translation dataset. Currently, the dataset contains approximately 60M entries and will continue to expand in the future.
BiST consists of two subsets, namely en-zh and zh-en, where the former represents the source language, collected from public data as real-world content; the latter represents the target… See the full description on the dataset page:
https://huggingface.co/datasets/yumatin/BiST.