The Dataset_JP-EN dataset comprises of three primary columns:
task: This column contains instruct prompts related to translate from Japanese text to English.
input: String in japanese.
expected_output: String in English.
The dataset is based on:
Helsinki-NLP/tatoeba_mt_train
facebook/flores
google/wmt24pp
Helsinki-NLP/opus-100