This repo contains an processed version of the Nunavut Hansard Inuktitut–English Parallel Corpus 3.0 dataset.
The dataset has been deduped and blank lines removed. Further, both corpora have lemma and part of speech information.
For English, this information has been added via SpaCy's en_core_web_trf model.
For Inuktitut, this information comes from automatic neural and rule-based morphological analyses of words courtesy of Jeffrey Micher.
Inuktitut additionally contains a translation of… See the full description on the dataset page:
https://huggingface.co/datasets/EdinburghNLP/nunavut-hansard-plusplus.