This dataset contains the "multi30k" dataset, which is the "task 1" dataset from here.
Each example consists of an "en" and a "de" feature. "en" is an English sentence, and "de" is the German translation of the English sentence.
Data Splits
The Multi30k dataset has 3 splits: train, validation, and test.
Dataset Split
Number of Instances in Split
Train
29,000
Validation
1,014
Test
1,000
Citation Information… See the full description on the dataset page: https://huggingface.co/datasets/bentrevett/multi30k.