The ssj500k training corpus contains about 500 000 tokens manually annotated on the levels of tokenisation,
sentence segmentation, morphosyntactic tagging, and lemmatisation. About half of the corpus is also manually annotated
with syntactic dependencies, named entities, and verbal multiword expressions. About a quarter of the corpus is also
annotated with semantic role labels. The morphosyntactic tags and syntactic dependencies are included both in the
JOS/MULTEXT-East framework, as well as in the framework of Universal Dependencies.