Release of the Mediomatix corpus, prepared by UZH and PHGR, to be used for research purposes only.
In mediomatix-raw, we release the full Mediomatix schoolbooks' text for each idiom. The data points are unaligned. See here for the mulit-parallel, aligned Mediomatix corpus.
We use the following Romansh idiom codes as subsets in the dataset:
Sursilvan: rm-sursilv
Sutsilvan: rm-sutsilv
Surmiran: rm-surmiran
Puter: rm-puter
Vallader: rm-vallader
The splits in… See the full description on the dataset page:
https://huggingface.co/datasets/ZurichNLP/mediomatix-raw.