Translation dataset with Vietnamese input prompt. This dataset was made with help from hieu1053/mtet There were 26 vietnamese prompts used to build the dataset. Below are the list of prompts and the prompt distributions for each split NOTE! I have filtered out rows['loss'] > 0.75 so this dataset is smaller than the original. Also, prompts such as Dịch câu sau "John is such a "funny guy"." are handled by code.
These were the… See the full description on the dataset page:
https://huggingface.co/datasets/nlplabtdtu/translation-text.