ASSETUK dataset, which was obtained by translating the ASSET dataset into Ukrainian and then manually checking the translation. The corpus is composed of 2000 validation and 359 test original sentences dataset/asset.{valid,test}.orig that were each simplified 10 times by different annotators dataset/asset.{valid.test}.simp.{0,1,2,3,4,5,6,7,8,9} (one sample per line).
If you have any questions, please contact the author:
Olha Kanishcheva (
kanichshevaolga@gmail.com)
If you use this dataset… See the full description on the dataset page:
https://huggingface.co/datasets/OKanishcheva/ASSETUK.