cleaned_PhoMT is a cleaned and quality-filtered subset of the PhoMT Vietnamese-English parallel corpus. The dataset is designed for machine translation research, particularly for training and evaluating Vietnamese ↔ English translation systems.
Starting from the original PhoMT corpus, we applied multiple data cleaning and quality control procedures to remove noisy, misaligned, and low-quality sentence pairs. The resulting dataset… See the full description on the dataset page: https://huggingface.co/datasets/Huyisbeee/cleaned_PhoMT.