This corpus is extracted from the JESC, with Japanese-English pairs.
For more information, see website below!
(https://nlp.stanford.edu/projects/jesc/index_ja.html)
JESC is the product of a collaboration between Stanford University, Google Brain, and Rakuten Institute of Technology. It was created by crawling the internet for movie and tv subtitles and aligining their captions. It is one of the largest freely available EN-JA corpus… See the full description on the dataset page: https://huggingface.co/datasets/nntsuzu/JESC.